1
00:00:00,240 --> 00:00:02,360
Hi, and welcome to the Neil
Ashton Podcast.

2
00:00:03,000 --> 00:00:05,880
In each episode, we explained
some of the fascinating ways

3
00:00:05,880 --> 00:00:09,080
that science and engineering are
changing the world around us.

4
00:00:09,760 --> 00:00:12,720
We talked to leading engineers
from elite level sports like

5
00:00:12,800 --> 00:00:16,840
cycling in Formula One to some
of the world's top academics to

6
00:00:16,840 --> 00:00:20,960
understand how fluid dynamics,
machine learning, supercomputing

7
00:00:21,360 --> 00:00:22,960
are bringing in a new era
discovery.

8
00:00:23,920 --> 00:00:27,040
We also hear some of their life
stories, their career advice,

9
00:00:27,560 --> 00:00:30,040
the lessons they've learned on
the way that I hope will be

10
00:00:30,040 --> 00:00:33,800
helpful to you too.
So sit back and enjoy this

11
00:00:33,800 --> 00:00:41,120
episode.
Hi, and welcome back to the Neil

12
00:00:41,120 --> 00:00:45,080
Ashton Podcast.
So today's episode is with

13
00:00:45,080 --> 00:00:47,840
somebody who I've actually
really enjoyed getting to know

14
00:00:47,840 --> 00:00:52,520
better and chatting it various
conferences that we've attended

15
00:00:52,520 --> 00:00:55,920
together and some mini symposia
that we've done, which is

16
00:00:56,160 --> 00:00:59,120
Johannes brand set up.
So he's got an interesting

17
00:00:59,120 --> 00:01:04,760
profile because not only is he a
professor at England's at EU

18
00:01:05,040 --> 00:01:09,960
Johannes Kepler University in
Austria, but he also founded or

19
00:01:10,040 --> 00:01:13,000
Co founded a startup recently
called MEAI.

20
00:01:14,520 --> 00:01:20,600
His background is one that is
very well positioned to help

21
00:01:20,600 --> 00:01:24,480
advance the state of machine
learning for, for for CFD and I

22
00:01:24,480 --> 00:01:27,400
guess CAE more broadly.
And yes, this is an episode on

23
00:01:27,400 --> 00:01:29,720
machine learning again, I had
sometimes feel bad that I keep

24
00:01:29,720 --> 00:01:35,240
doing this, but I think this
season there's a mixture of of

25
00:01:35,240 --> 00:01:38,920
episodes and I still feel that
there is value in going through

26
00:01:38,920 --> 00:01:41,640
the machine learning topic.
And it's selfishly something

27
00:01:41,640 --> 00:01:44,560
that I I'm constantly interested
in learning more about and

28
00:01:44,560 --> 00:01:46,880
learning more about.
It is one thing that I always do

29
00:01:46,880 --> 00:01:50,200
when I speak to Yanis.
And he has an interesting

30
00:01:50,200 --> 00:01:55,760
background because whilst he has
AI guess a high energy physics

31
00:01:55,800 --> 00:02:01,640
background that he he, he did
and he was at CERN and he worked

32
00:02:01,640 --> 00:02:04,840
in that sort of area.
He then moved into the machine

33
00:02:04,840 --> 00:02:10,400
learning side, I guess, having
spent time with Max Welling at

34
00:02:10,400 --> 00:02:13,040
the Amsterdam University in the
machine learning lab.

35
00:02:13,320 --> 00:02:17,360
And then I think he also, well,
I don't think I know that he was

36
00:02:17,360 --> 00:02:19,720
then at Microsoft Research where
Max was also out.

37
00:02:19,720 --> 00:02:21,720
And so the two of them
collaborated a lot.

38
00:02:22,400 --> 00:02:25,280
You know, if you don't know who
Max Welling is, well, first

39
00:02:25,280 --> 00:02:27,520
Google, his name is one of The
Pioneers of machine learning.

40
00:02:27,520 --> 00:02:30,520
But I also did an episode with
him in the last season that I

41
00:02:30,520 --> 00:02:34,720
think was really interesting and
I think he must have been

42
00:02:34,720 --> 00:02:36,840
inspired by Max a little bit,
even though we didn't talk about

43
00:02:36,840 --> 00:02:40,280
it in this episode, because Max
also has been a serial start up

44
00:02:40,720 --> 00:02:44,800
founder.
And the reason I said that he

45
00:02:44,800 --> 00:02:47,040
has an interesting background is
because whilst he was at

46
00:02:47,040 --> 00:02:51,560
Microsoft, he also worked on the
Aurora machine learning for web

47
00:02:51,560 --> 00:02:55,840
and climate project.
That I think is extremely useful

48
00:02:55,840 --> 00:02:58,200
when you have people who have
gone through these major

49
00:02:58,960 --> 00:03:01,880
projects in another field
because they bring with them

50
00:03:01,880 --> 00:03:04,080
lessons and learnings that are
important.

51
00:03:04,440 --> 00:03:09,400
And, and what Johannes has been
doing is first of all, bringing

52
00:03:09,400 --> 00:03:13,240
a very academic mindset to this
academic in the terms of

53
00:03:13,360 --> 00:03:16,560
publishing and transparency,
which I think is very welcome.

54
00:03:16,560 --> 00:03:18,880
So you'll find in the link in
the YouTube a couple of the

55
00:03:18,880 --> 00:03:24,200
papers that he's published and
and his startup has published a

56
00:03:24,200 --> 00:03:27,480
new sort of transformer based
model that I, at least from my

57
00:03:27,480 --> 00:03:32,760
reading, is unique and certainly
seems to be at the one of the

58
00:03:32,760 --> 00:03:34,720
most cutting edge in
state-of-the-art models out

59
00:03:34,720 --> 00:03:38,160
there today, both in terms of
conceptual but also the accuracy

60
00:03:38,160 --> 00:03:40,200
on the data sets that they've
shown.

61
00:03:40,880 --> 00:03:45,280
But I admire his vision and his
willingness to try and solve the

62
00:03:45,280 --> 00:03:48,840
problem rather than being
focused on the model, as in some

63
00:03:48,840 --> 00:03:53,080
people seem to, once they come
up with a model, fixate on that

64
00:03:53,080 --> 00:03:57,080
being it, rather than being
willing to consider that their

65
00:03:57,080 --> 00:03:59,920
model may be the right thing at
the time, but then there'll be

66
00:03:59,920 --> 00:04:02,800
other models that get better.
And he, his willingness to

67
00:04:02,800 --> 00:04:06,800
accept that, I think is a breath
of fresh air and definitely will

68
00:04:06,800 --> 00:04:09,880
help the community to, to
involve, to evolve.

69
00:04:10,800 --> 00:04:12,240
So that's what we talked about
today.

70
00:04:12,240 --> 00:04:15,560
Really we, we tried to go
through and discuss, you know,

71
00:04:15,680 --> 00:04:18,720
more the general topics around
machine learning, but really

72
00:04:18,720 --> 00:04:22,720
diving into this transformer
based approach that he has.

73
00:04:22,720 --> 00:04:25,320
And I'm trying to understand
some of his thoughts around the

74
00:04:25,320 --> 00:04:28,280
similarities to null operators,
some of the slight differences,

75
00:04:29,440 --> 00:04:32,400
some of the links to I guess
graph neural Nets and mesh graph

76
00:04:32,400 --> 00:04:34,480
Nets that probably most people
are aware of.

77
00:04:35,080 --> 00:04:38,160
And then we talked a little bit
about the process of forming the

78
00:04:38,160 --> 00:04:41,760
startup and some discussions
around future topics.

79
00:04:43,160 --> 00:04:45,760
As with anything, I always feel
bad because I finished the

80
00:04:45,760 --> 00:04:49,000
episode and I think I really
should have discussed more or,

81
00:04:49,040 --> 00:04:52,120
or, or dived into certain
topics, but I'm conscious it was

82
00:04:52,120 --> 00:04:56,280
already over, you know, an hour
20 and people will, you know,

83
00:04:56,280 --> 00:04:58,440
fall asleep listening to some of
these episodes.

84
00:04:58,840 --> 00:05:02,400
But hopefully it's enough to
spark your interest and

85
00:05:02,400 --> 00:05:05,200
encourage you to read more about
the topics that we discussed.

86
00:05:05,720 --> 00:05:08,400
And please leave any comments or
let me know if you have any

87
00:05:08,400 --> 00:05:10,720
suggestions for for other
topics.

88
00:05:11,160 --> 00:05:15,000
But for now, please sit back and
enjoy what I found an extremely

89
00:05:15,000 --> 00:05:17,480
interesting episode with
Johannes Bransetta.

90
00:05:18,720 --> 00:05:22,440
Cool.
All right, thanks for joining me

91
00:05:22,440 --> 00:05:24,760
today.
I wanted to have a chat with you

92
00:05:24,760 --> 00:05:27,960
for a while.
We've had some nice discussions

93
00:05:27,960 --> 00:05:30,680
over wine and and coffee, but
nothing public.

94
00:05:30,680 --> 00:05:36,320
So this is the public version of
our discussions, but maybe it'd

95
00:05:36,320 --> 00:05:39,800
be good for people to learn a
little bit more about you.

96
00:05:39,800 --> 00:05:42,880
You know, where you're at today.
Yeah.

97
00:05:43,440 --> 00:05:44,600
Who?
Who is Johannes?

98
00:05:45,520 --> 00:05:47,960
Yeah, thanks first of all.
Thanks really Neil for for

99
00:05:47,960 --> 00:05:49,520
having me.
It's it's a great pleasure

100
00:05:49,520 --> 00:05:53,680
finally seeing this guitars here
in a one to one setting, so to

101
00:05:53,680 --> 00:05:56,800
say.
Yeah, my background is actually

102
00:05:56,800 --> 00:06:00,560
I'm I'm a learned physicist,
some say a failed physicist.

103
00:06:01,000 --> 00:06:05,080
But after my PHDI switched to
machine learning, as many people

104
00:06:05,280 --> 00:06:07,560
also did.
I have spent time with

105
00:06:07,560 --> 00:06:11,320
Sepulchreiter and three years in
Amsterdam with Maxwelling.

106
00:06:11,320 --> 00:06:13,880
I was in industry for two years
at Microsoft Research.

107
00:06:15,080 --> 00:06:19,360
And there at Microsoft Research
I really, really discovered my

108
00:06:19,400 --> 00:06:21,920
likings for large scale
simulation, especially for

109
00:06:21,920 --> 00:06:24,320
weather and climate modelling,
because these are like some of

110
00:06:24,320 --> 00:06:26,480
the biggest problems you can
have and you can tackle with

111
00:06:26,480 --> 00:06:29,240
machine learning.
I also saw the transformative

112
00:06:29,240 --> 00:06:35,960
impact on this, on the systems,
and that was then the point

113
00:06:35,960 --> 00:06:40,400
where I decided to to do my own
thing to, but to apply it not to

114
00:06:40,760 --> 00:06:44,960
known problems like weather, but
rather uncharted territories

115
00:06:44,960 --> 00:06:47,520
like engineering, simulation and
this kind of thing.

116
00:06:48,600 --> 00:06:52,320
And I decided to get my own
group at university to build up

117
00:06:52,320 --> 00:06:56,360
my knowledge base back in
Austria, where I'm coming from.

118
00:06:56,880 --> 00:06:59,640
And along the road.
It also happened that I founded

119
00:06:59,640 --> 00:07:03,320
a start up MEII where we do this
simulations at scale and where

120
00:07:03,320 --> 00:07:06,520
things are coming together with
a bit more compute and a bit

121
00:07:06,520 --> 00:07:10,360
more resources.
Nice.

122
00:07:10,360 --> 00:07:17,160
Well, we'll we need to dive into
each of those, but maybe the

123
00:07:17,160 --> 00:07:21,440
first one was the Aurora
project.

124
00:07:21,440 --> 00:07:27,000
I guess this is maybe it, it was
subtaneously both a motivation,

125
00:07:27,000 --> 00:07:29,520
I think certainly for me and
others who saw what was

126
00:07:29,520 --> 00:07:32,040
happening in weather and climate
and was saying, you know, surely

127
00:07:32,040 --> 00:07:35,840
we could do that in engineering.
But also it was an example of a,

128
00:07:37,120 --> 00:07:39,320
yeah, a real large scale
project.

129
00:07:40,800 --> 00:07:42,960
So what was it like working on
that Aurora?

130
00:07:42,960 --> 00:07:46,320
What was the key lessons you
learned both from maybe an ML

131
00:07:46,320 --> 00:07:52,040
point of view, but also a
project and a team point of

132
00:07:52,040 --> 00:07:55,360
view?
So the actual Aurora team was

133
00:07:55,360 --> 00:07:57,600
pretty small.
I mean, all the credits to to

134
00:07:57,640 --> 00:08:02,360
Chris Bodner versus Pransma,
Megan Stanley and a Lucic who

135
00:08:02,360 --> 00:08:05,840
did that, all the heavy lifting
afterwards to finalize this

136
00:08:05,840 --> 00:08:07,720
project, to train to build the
data loader.

137
00:08:09,640 --> 00:08:13,360
The Aurora was the learnings I
think are threefold.

138
00:08:13,440 --> 00:08:17,160
First of all, with computer
vision tricks, you can get very,

139
00:08:17,160 --> 00:08:21,160
very far in engineering or
scientific application.

140
00:08:21,160 --> 00:08:24,880
In the end it was a swing
transformer trained in in a 3D

141
00:08:25,040 --> 00:08:28,560
swing transformer.
Secondly that the the data

142
00:08:28,680 --> 00:08:35,039
engineering is potentially the
hardest one and and thirdly that

143
00:08:35,039 --> 00:08:39,039
weather has a very unfair
advantage for machine learning

144
00:08:39,039 --> 00:08:41,480
because nobody knows how weather
is functioning.

145
00:08:41,880 --> 00:08:47,480
So you, you will always train a
machine learning model on like

146
00:08:47,480 --> 00:08:51,520
actual data, which basically has
the actual physical laws hidden.

147
00:08:51,880 --> 00:08:57,320
Whereas the, the the numerical
methods, they have to come up

148
00:08:57,320 --> 00:09:02,120
with these laws by themselves.
And that already brings these,

149
00:09:02,120 --> 00:09:07,600
these key interests of mine,
which is that that neural

150
00:09:07,600 --> 00:09:09,960
surrogates will never replace
numerics.

151
00:09:09,960 --> 00:09:13,440
They will just be a different
branch which depending on how

152
00:09:13,440 --> 00:09:15,840
you play the card acts in your
favour.

153
00:09:16,320 --> 00:09:19,200
And that it's all about
engineering and scaling these,

154
00:09:19,200 --> 00:09:24,000
these things to this, to this
basically to this problems where

155
00:09:24,000 --> 00:09:26,480
they really have an impact and
there really matters.

156
00:09:28,440 --> 00:09:31,360
And the impact for Aurora, you
could see if you look through

157
00:09:31,360 --> 00:09:35,640
this, this small, through this
downstream tasks that the larger

158
00:09:35,640 --> 00:09:38,160
these challenges are, the more
impact you can basically

159
00:09:38,160 --> 00:09:42,480
generate.
Yeah, it seemed to definitely

160
00:09:43,240 --> 00:09:44,920
spur on.
And now it seems that there's

161
00:09:44,920 --> 00:09:49,040
been so many of the weather and
climate models, it's almost got

162
00:09:49,320 --> 00:09:52,240
not congested, but there's
there's, there's quite a few

163
00:09:52,280 --> 00:09:56,760
coming out.
And so it's not obviously so

164
00:09:56,760 --> 00:10:01,640
clear how each are progressing
past each other and and how much

165
00:10:01,640 --> 00:10:06,280
are those of just an individual
groups need to have their own

166
00:10:06,280 --> 00:10:11,320
model, if you know what I mean.
But did you, I guess one of the

167
00:10:11,320 --> 00:10:14,080
big things that I want to
discuss today with you in

168
00:10:14,080 --> 00:10:22,160
particular was to dive into, I
guess where machine learning,

169
00:10:24,680 --> 00:10:28,120
how can I phrase this machine
learning for engineering, for

170
00:10:28,120 --> 00:10:31,080
CFD, for CAE.
You know, it's something I've

171
00:10:31,080 --> 00:10:36,520
spoken to a few people about and
there's a lot going on, but it's

172
00:10:37,120 --> 00:10:41,720
hasn't always been so clear
where the big breakthroughs will

173
00:10:41,720 --> 00:10:46,800
be or how, how good are we today
and where we're going.

174
00:10:46,800 --> 00:10:51,640
It it's it's still seems a
congested space with different

175
00:10:51,640 --> 00:10:54,920
opinions, you know, with pins on
one side and then these sort of

176
00:10:54,920 --> 00:10:56,640
mesh graph Nets and other
things.

177
00:10:56,640 --> 00:11:00,720
So it's a bit of a, an
open-ended question to you, but

178
00:11:00,720 --> 00:11:04,800
I guess like where do you see
the state-of-the-art?

179
00:11:04,800 --> 00:11:09,400
What what's been your journey
from the sort of mesh graph Nets

180
00:11:09,800 --> 00:11:15,120
towards your current thinking?
Like how, how have you try to

181
00:11:15,120 --> 00:11:18,000
solve this problem, I guess.
And let's, let's say, let's take

182
00:11:18,000 --> 00:11:22,240
CFD, maybe is the the the
problem like automotive

183
00:11:22,360 --> 00:11:25,920
aerodynamics for example?
Yeah, that's a very good, my

184
00:11:25,920 --> 00:11:29,280
favorite topic actually making
the connections to weather.

185
00:11:29,280 --> 00:11:32,280
Let's let's say what what
basically was driving this

186
00:11:32,280 --> 00:11:35,840
weather modelling.
I would say that NVIDIA was the

187
00:11:35,840 --> 00:11:40,040
first with the forecast net
paper to to bring out a model

188
00:11:40,040 --> 00:11:43,200
which worked that was trained on
error 5, which is a publicly

189
00:11:43,200 --> 00:11:47,120
available large scale data set.
Depending on how you sample you

190
00:11:47,120 --> 00:11:51,520
get roughly a petabyte of data
or a few 100 terabytes of data

191
00:11:51,520 --> 00:11:54,720
from it.
And then you do basically mean

192
00:11:54,720 --> 00:11:59,680
squared error training that that
was the the scene set and and as

193
00:11:59,680 --> 00:12:04,280
soon as this problem was defined
of this input output relation of

194
00:12:04,280 --> 00:12:08,280
this metrics of of what to test,
then people did what, what they

195
00:12:08,280 --> 00:12:11,760
do best.
They optimize on this sort of

196
00:12:11,760 --> 00:12:15,120
problems.
If you go to engineering, things

197
00:12:15,120 --> 00:12:17,440
are very different in, in many,
many aspects.

198
00:12:17,440 --> 00:12:19,920
So first of all, we don't have
this data set.

199
00:12:19,920 --> 00:12:21,760
We don't have an error five data
set.

200
00:12:22,200 --> 00:12:25,040
And even if you had an error
five data set, I mean, there are

201
00:12:25,040 --> 00:12:30,920
now at least 2-3 publicly
available CFD data sets.

202
00:12:30,920 --> 00:12:35,600
Needless to say that my favorite
one is the Tri ML data set which

203
00:12:35,600 --> 00:12:40,120
are industrial standard.
But even if these data sets are

204
00:12:40,120 --> 00:12:43,040
out there, people don't know how
to train on them.

205
00:12:43,040 --> 00:12:47,040
And everyone trains differently.
And this comes from the fact

206
00:12:47,040 --> 00:12:49,880
that we just don't know what you
really want to optimize.

207
00:12:50,240 --> 00:12:53,680
There is people who sub sample
parts of the surface and map to,

208
00:12:53,680 --> 00:12:56,160
to pressure values on the
surface, people who sub sample

209
00:12:56,160 --> 00:12:58,840
part of surface and volume and
so on and so forth.

210
00:12:59,160 --> 00:13:02,800
So similarly to whether we have
to figure out what's the right

211
00:13:02,920 --> 00:13:08,040
task, what's the right learning
task, which we which we have to

212
00:13:08,160 --> 00:13:12,320
to do or like the tasks in, in
order to really develop the

213
00:13:12,320 --> 00:13:15,680
right models in the right
frameworks for that.

214
00:13:17,800 --> 00:13:23,720
Secondly, I think what is also
very different to weather is

215
00:13:23,720 --> 00:13:25,880
that we have this input output
relation.

216
00:13:25,880 --> 00:13:30,440
So in weather, it's always clear
you take time T and you map to T

217
00:13:30,440 --> 00:13:33,480
+ 1.
That's basically a segmentation

218
00:13:33,480 --> 00:13:37,960
task on steroids because in the
end every pixel on the Earth

219
00:13:37,960 --> 00:13:43,120
gets mapped to but with pixel on
the earth like a time time point

220
00:13:43,120 --> 00:13:48,080
later, which you can basically
you can take all the models from

221
00:13:48,080 --> 00:13:51,080
computer vision.
And obviously there is some

222
00:13:51,080 --> 00:13:54,160
tricks with resolution and and
if you take the sphere into

223
00:13:54,160 --> 00:13:57,600
account, yes or not.
But in the end it's it's a pixel

224
00:13:58,240 --> 00:14:01,520
segmentation task.
And this is definitely not true

225
00:14:01,520 --> 00:14:08,200
for for many CFD related tasks
because the task there is from a

226
00:14:08,200 --> 00:14:12,600
pure numerical perspective, you
have a geometry which you can

227
00:14:12,600 --> 00:14:15,720
characterize with couple of
parameters, probably less than

228
00:14:15,720 --> 00:14:17,600
100 even.
And then you get this

229
00:14:17,600 --> 00:14:21,040
full-fledged flow fields around
and on the car so that the

230
00:14:21,040 --> 00:14:24,440
actual input to output ratio is
very, very different.

231
00:14:25,760 --> 00:14:30,560
And we don't have architectures,
we don't have frameworks, we

232
00:14:30,560 --> 00:14:33,120
don't have an understanding of
how to tackle that.

233
00:14:34,080 --> 00:14:36,640
But that makes it so exciting I
would say.

234
00:14:37,040 --> 00:14:42,000
And maybe thirdly, if you, if
you really look at these weather

235
00:14:42,000 --> 00:14:44,960
models and what they can do with
hurricane predictions and so on,

236
00:14:44,960 --> 00:14:48,080
obviously there's always room
for improvement and obviously

237
00:14:48,080 --> 00:14:50,040
you can go farther and farther.
But if you just look what

238
00:14:50,040 --> 00:14:53,160
happened in the last two years
and the resources and weather

239
00:14:53,160 --> 00:14:56,840
are I would say it's tiny, tiny
fractions to the resources in

240
00:14:56,840 --> 00:15:00,600
LLM and and video generation.
You see this tremendous progress

241
00:15:00,600 --> 00:15:03,880
and what what hard problems and
hard multi scale problems we are

242
00:15:03,880 --> 00:15:06,520
already able to model and which
we can able better than

243
00:15:06,520 --> 00:15:09,080
numerics.
So it's very clear that if done

244
00:15:09,080 --> 00:15:11,880
rightly and if the right
incentives are there, this

245
00:15:12,000 --> 00:15:15,680
progress and this type of
complexity we're able to do in

246
00:15:15,680 --> 00:15:18,480
in CFD.
Obviously then we have to play

247
00:15:18,480 --> 00:15:21,080
the game numerics and ML
together.

248
00:15:22,120 --> 00:15:24,920
Yeah, which is then I think the
hardest task to answer.

249
00:15:27,200 --> 00:15:30,480
Yeah, that, that does seem, I
guess the joke that the Earth is

250
00:15:30,480 --> 00:15:33,000
always the Earth.
And so the geometry is sort of,

251
00:15:33,520 --> 00:15:35,320
you know, not so much the
problem.

252
00:15:36,880 --> 00:15:42,440
And yeah, I guess it all depends
on what is the data.

253
00:15:44,320 --> 00:15:48,400
You're right, in the CFD
example, the geometries could be

254
00:15:48,400 --> 00:15:52,160
hugely varying and that link
between the geometry and the

255
00:15:52,160 --> 00:15:56,960
volume is a Yeah, it it's so the
geometry is geometry and the

256
00:15:56,960 --> 00:15:59,160
boundary conditions is
ultimately what affects the

257
00:15:59,160 --> 00:16:03,320
flows.
So where have you seen, I guess

258
00:16:03,320 --> 00:16:06,280
a lot of people listening to
this and myself included,

259
00:16:06,880 --> 00:16:10,840
probably had their first
introduction with, you know,

260
00:16:10,840 --> 00:16:14,440
what the Deep Mind team did with
with mesh graph Nets.

261
00:16:14,600 --> 00:16:17,480
I mean, there was obviously
stuff before that, but that felt

262
00:16:17,480 --> 00:16:24,120
like the bit when everyone stood
up and and noticed why?

263
00:16:24,920 --> 00:16:28,720
Why do you not feel that that is
the right approach?

264
00:16:30,800 --> 00:16:34,440
It's always right or wrong.
It's always hard to to to judge.

265
00:16:34,440 --> 00:16:39,480
But I think that the, the mesh
craft net and, and, and and and

266
00:16:39,480 --> 00:16:42,720
graph neural network simulators
from the Peter Battaglia group

267
00:16:42,720 --> 00:16:47,160
were tremendously important
milestone in this field, mostly

268
00:16:47,160 --> 00:16:49,760
because they, they were writing
down this learning problem.

269
00:16:50,600 --> 00:16:55,720
And I myself went into this
field of simulations because of

270
00:16:55,720 --> 00:16:58,640
these papers trying to model
particle particle systems.

271
00:17:00,240 --> 00:17:03,480
What, what we noticed back in
the days, I mean, this is like

272
00:17:03,480 --> 00:17:06,680
5-6 years ago, that it's
tremendously hard to model all

273
00:17:06,680 --> 00:17:09,319
particle particle interaction as
numerics is doing.

274
00:17:09,319 --> 00:17:13,440
So if you have this 10,000
system of particles and they

275
00:17:13,440 --> 00:17:15,520
interact and you have you have
something like a graph neural

276
00:17:15,520 --> 00:17:18,359
network and you have to make all
these interactions correctly

277
00:17:18,359 --> 00:17:22,960
because otherwise the temporal
integration will will just,

278
00:17:24,200 --> 00:17:29,520
yeah, just go, go, go crazy.
That's just a very hard learning

279
00:17:29,520 --> 00:17:30,720
problem.
And this is not what machine

280
00:17:30,720 --> 00:17:32,640
learning is really good at.
Machine learning is good at to

281
00:17:32,640 --> 00:17:35,920
understand the global dynamics,
to understand the behavior of

282
00:17:35,920 --> 00:17:38,960
the system, to understand where
things are going on a global

283
00:17:38,960 --> 00:17:41,480
scale, but not on a particle
particle scale.

284
00:17:42,840 --> 00:17:46,760
So I, I think from, from, from
setting up this learning

285
00:17:46,760 --> 00:17:49,120
problem, this was a tremendously
important step.

286
00:17:49,360 --> 00:17:54,760
Obviously new ways of, of, of of
models are coming the way for me

287
00:17:54,760 --> 00:17:57,480
personally, it was a big
breakthrough that we started to,

288
00:17:57,480 --> 00:18:01,320
to see things as fields.
So everything in my, in my head

289
00:18:01,320 --> 00:18:04,560
is a field.
And because you a field is

290
00:18:04,560 --> 00:18:06,120
something which evolves over
time.

291
00:18:06,440 --> 00:18:09,160
It can be an occupancy field,
which tells you where is mass,

292
00:18:09,160 --> 00:18:11,480
where is no mass.
It can be assigned distance

293
00:18:11,480 --> 00:18:13,480
field.
It can be a velocity field, it

294
00:18:13,480 --> 00:18:15,600
can be a displacement field,
whatever field you want.

295
00:18:16,240 --> 00:18:21,720
But it's much easier for a, for
a model to, to tackle, which is

296
00:18:21,720 --> 00:18:24,840
actually obviously the, the next
step if you think of this whole

297
00:18:24,840 --> 00:18:28,320
scientific machine learning and
no operator and, and, and so on

298
00:18:28,320 --> 00:18:31,280
and so forth.
And in the end, I mean, this is

299
00:18:31,280 --> 00:18:33,480
also how you, you model the,
the, the weather.

300
00:18:33,480 --> 00:18:36,440
So I, I think conceptually from
me personally, from me, I don't

301
00:18:36,440 --> 00:18:39,760
know how it is for others.
Understanding that you can model

302
00:18:39,760 --> 00:18:42,920
large scale systems as a field
was, was breaking a lot of

303
00:18:42,920 --> 00:18:45,640
boundaries in, in terms of
scalability.

304
00:18:45,640 --> 00:18:50,200
Because if you model a system of
10 million, 100 million

305
00:18:50,200 --> 00:18:53,520
particles as a field, it's
computation and not the problem.

306
00:18:53,520 --> 00:18:55,280
It's from a learning
perspective, it's not the

307
00:18:55,280 --> 00:18:57,600
problem because you just don't
need all these particles.

308
00:18:57,600 --> 00:19:01,520
You you'll learn the dynamics
with the very heavily subsampled

309
00:19:03,360 --> 00:19:08,400
representation and and and and
obtaining the full field is also

310
00:19:08,400 --> 00:19:11,040
not a problem because this this
exists.

311
00:19:11,440 --> 00:19:15,280
So I would say for me
understanding what machine

312
00:19:15,280 --> 00:19:18,480
learning is really good at and
and and and and and what

313
00:19:18,680 --> 00:19:22,800
numerics is really good at and
and and and and playing those

314
00:19:22,800 --> 00:19:27,800
two on, on on those two fronts
was was was making the the

315
00:19:27,800 --> 00:19:29,440
difference.
Yeah.

316
00:19:30,800 --> 00:19:36,280
Well, one thing that I've maybe
would appreciate you explaining

317
00:19:36,560 --> 00:19:41,800
to the audience and maybe let's
first introduce some of the work

318
00:19:41,800 --> 00:19:45,160
that you've done in this recent
paper of yours.

319
00:19:46,600 --> 00:19:51,160
So you you developed this
transformer based approach, what

320
00:19:51,160 --> 00:19:53,920
you call it anchor based
universal physics.

321
00:19:53,920 --> 00:19:54,800
Trump What?
What?

322
00:19:54,880 --> 00:19:58,800
Well, maybe you explain what?
Yeah, we.

323
00:19:59,440 --> 00:20:03,680
Have actually a couple of works
of 1 is which we call Universal

324
00:20:03,680 --> 00:20:05,960
Physics transformer.
That's basically how we do

325
00:20:05,960 --> 00:20:08,200
latent space modelling of
logical system.

326
00:20:08,600 --> 00:20:11,080
Then we did something which is
called Neural DM where we

327
00:20:11,080 --> 00:20:14,640
applied those to to multi flows
and and and multi.

328
00:20:15,200 --> 00:20:19,960
Particle systems and then we,
we, we recently applied this to

329
00:20:19,960 --> 00:20:23,880
to large scale CFD similar
framework, similar approach,

330
00:20:24,680 --> 00:20:27,520
which we call anchored branched
because these are the two

331
00:20:27,520 --> 00:20:31,160
ingredients which which which we
need to do in order to get this

332
00:20:31,160 --> 00:20:37,360
to work on large systems.
So in that context, what I'm

333
00:20:37,360 --> 00:20:40,960
hoping you could do is maybe
explain a little bit, because

334
00:20:40,960 --> 00:20:42,920
there's a lot of people who
listen to this podcast who are

335
00:20:42,920 --> 00:20:48,440
maybe not ML specialists, but
are CFD primarily specialists

336
00:20:48,440 --> 00:20:52,280
who are now, you know, more and
more interested in ML and maybe

337
00:20:52,280 --> 00:20:55,960
actually developing it and, and
yeah, sort of progressing.

338
00:20:56,320 --> 00:21:03,520
And you mentioned a really
interesting point to me and I'm

339
00:21:03,520 --> 00:21:06,600
hoping you can explain this in a
way that is understandable by

340
00:21:06,600 --> 00:21:11,080
others is that I, and I guess
some of this is a maths

341
00:21:11,080 --> 00:21:14,200
exercise, but you know, and I've
been guilty of this.

342
00:21:14,200 --> 00:21:17,640
Sometimes there's this
hierarchical thought process of

343
00:21:17,640 --> 00:21:21,360
being like, OK, you've got some
sort of graph based thing where

344
00:21:21,360 --> 00:21:24,320
it's more like point to point.
Then you've got a new operator

345
00:21:24,320 --> 00:21:27,600
where it's like more to
solutions or fields.

346
00:21:28,240 --> 00:21:31,480
And then I initially thought of,
OK, and then you have

347
00:21:31,480 --> 00:21:34,440
Transformers.
And yet you see some people who

348
00:21:34,440 --> 00:21:37,120
say, well, everything is a new
operator if you do the maths or

349
00:21:37,120 --> 00:21:41,720
design it.
So can you explain and maybe

350
00:21:41,720 --> 00:21:45,280
dive a little bit deeper into
that concept of solution mapping

351
00:21:45,280 --> 00:21:47,880
or field mapping and then how it
lakes transforms?

352
00:21:47,880 --> 00:21:50,960
Because I think it is a little
bit confusing for for for

353
00:21:50,960 --> 00:21:53,160
people.
So if yeah, if we could maybe

354
00:21:53,160 --> 00:21:58,080
get into that, that would be in
in the context of let's say this

355
00:21:58,280 --> 00:22:00,680
drive air, ML or car, however
you want to explain it.

356
00:22:01,280 --> 00:22:04,200
Yeah, happily to do so.
So for me, in fact neural

357
00:22:04,200 --> 00:22:06,760
operate and this is also how it
is introduced is something which

358
00:22:06,760 --> 00:22:12,000
maps between function spaces,
which is very yeah, which is not

359
00:22:12,000 --> 00:22:15,720
very informative if you if you,
because what does it mean?

360
00:22:15,720 --> 00:22:19,960
You map between function spaces.
But I always think of it as you

361
00:22:19,960 --> 00:22:22,760
have two functions, one input
function, 1 output function.

362
00:22:23,000 --> 00:22:28,040
And no matter what you sample
from the input function and what

363
00:22:28,040 --> 00:22:30,240
you predict on the output
function, this has to be

364
00:22:30,240 --> 00:22:32,960
fulfilled.
Meaning that if a sample 10

365
00:22:32,960 --> 00:22:36,760
points from the input function,
I should be able to predict as

366
00:22:36,760 --> 00:22:39,680
many points as I want from the
output function and those points

367
00:22:39,680 --> 00:22:41,600
I predict are actually on the
output function.

368
00:22:41,960 --> 00:22:46,760
If a sample with the same
network with the same operator 6

369
00:22:46,760 --> 00:22:50,160
points from the input function,
I should also get the same 10

370
00:22:50,160 --> 00:22:53,640
points and if I want more
points, I get more points on the

371
00:22:53,640 --> 00:22:56,400
output function.
Obviously there is a limit.

372
00:22:56,400 --> 00:22:59,800
If a sample to little input
points, then obviously the

373
00:22:59,800 --> 00:23:02,600
approximation network is not
able to really get the

374
00:23:02,600 --> 00:23:06,920
information needs.
But if if this this this

375
00:23:06,920 --> 00:23:10,360
sampling limit is is reached no
matter how many points and where

376
00:23:10,360 --> 00:23:13,040
those points are spaced, it
should always give you some sort

377
00:23:13,040 --> 00:23:16,160
of representation which allow
you to construct the output

378
00:23:16,160 --> 00:23:17,720
function.
So in the latent space this

379
00:23:17,720 --> 00:23:23,360
mapping is, then is then, so to
say, resolution agnostic and and

380
00:23:23,360 --> 00:23:28,120
that with some small tricks you
can do with convolutions, with

381
00:23:28,120 --> 00:23:31,040
full node operators, with
Transformers and so on and so

382
00:23:31,040 --> 00:23:33,280
forth.
And obviously Transformers with

383
00:23:33,280 --> 00:23:36,880
their.
Can I if I've seen to it just

384
00:23:36,880 --> 00:23:40,480
for one second, just to make it
even more clearer, when you say

385
00:23:41,160 --> 00:23:44,680
sampling the input function and
the output function by input

386
00:23:44,680 --> 00:23:49,480
function, you're meaning like
the geometry surface in this

387
00:23:49,480 --> 00:23:54,000
conceptual sense and the output
function being let's say the

388
00:23:54,000 --> 00:23:56,600
volume.
Would that be 1 interpretation?

389
00:23:56,840 --> 00:23:59,600
Yes, for example you can.
You can think of it that the

390
00:23:59,600 --> 00:24:03,040
input is, the is the geometry
and the output is either the

391
00:24:03,040 --> 00:24:07,120
field on the geometry or the
field in the volume, or both

392
00:24:07,440 --> 00:24:11,400
actually.
And the idea of this so where

393
00:24:11,480 --> 00:24:12,600
you know, this sounds like a bit
of AI.

394
00:24:12,600 --> 00:24:16,560
Remember when this first came
out, I was I think it was billed

395
00:24:16,600 --> 00:24:22,240
as a resolution independent.
So it which always to ACFD

396
00:24:22,240 --> 00:24:27,160
person feels like the no free
lunch sort of theorem.

397
00:24:27,160 --> 00:24:30,200
Like how can this possibly Yeah,
you know, how can that be true?

398
00:24:30,200 --> 00:24:33,800
You're basically saying I don't
need to map the entire geometry,

399
00:24:33,800 --> 00:24:36,280
I can just take a few and still
get the same answers.

400
00:24:36,280 --> 00:24:41,880
So where's the catch?
I I guess on this?

401
00:24:42,760 --> 00:24:44,120
Yeah.
I mean that that's a very good

402
00:24:44,120 --> 00:24:45,960
point.
And I think I can make my, my

403
00:24:45,960 --> 00:24:49,880
point very clear with, with
really now going to CFD and,

404
00:24:49,880 --> 00:24:53,080
and, and before that, I really
have to say, if people think of

405
00:24:53,080 --> 00:24:57,720
resolution, we always think
again of this segmentation as,

406
00:24:57,760 --> 00:25:00,920
as we do with weather, right?
You have a certain number of

407
00:25:00,920 --> 00:25:04,120
longitude and latitude points.
And if we double those points,

408
00:25:04,560 --> 00:25:08,120
do we get higher resolution or
do we just have an interpolation

409
00:25:08,120 --> 00:25:11,560
and so on and so forth.
So can we just run our whatever

410
00:25:11,560 --> 00:25:14,480
unit and interpolate between
this resolution or is there a

411
00:25:14,480 --> 00:25:17,760
way to get this stuff resolved?
So there's always this input

412
00:25:17,800 --> 00:25:23,200
output mapping and obviously it
will end up with interpolation

413
00:25:23,200 --> 00:25:24,680
effect.
So either you interpolate

414
00:25:24,680 --> 00:25:27,320
somewhere in your network or you
interpolate on the output grid

415
00:25:27,320 --> 00:25:30,040
because as you said, no free
lunch.

416
00:25:30,560 --> 00:25:34,840
But if we go to CFD, things are
a bit different, especially that

417
00:25:34,840 --> 00:25:37,600
the input output resolution, as
I said before, is very different

418
00:25:37,600 --> 00:25:42,360
or input output connection.
And, and actually we show in

419
00:25:42,360 --> 00:25:47,120
this recent paper that if you
subsample a few of the points,

420
00:25:47,120 --> 00:25:50,600
so both on geometry and surface
and, and you do some, some

421
00:25:50,600 --> 00:25:53,280
training, there is a certain
amount of points, but the

422
00:25:53,280 --> 00:25:56,760
performance stagnates.
So if it's, it's for for tribe

423
00:25:56,760 --> 00:26:04,480
ML, this is roughly 128 Ki think
or 256 I, I, I don't know by

424
00:26:04,480 --> 00:26:07,560
heart.
And if you then add more points,

425
00:26:07,560 --> 00:26:10,320
the performance is not getting
better because all the, the

426
00:26:10,320 --> 00:26:13,200
information which in neural
network needs is in those

427
00:26:13,200 --> 00:26:16,680
points.
And, and if you sample it, let's

428
00:26:16,680 --> 00:26:20,600
say cleverly so, so that the
distribution really covers the

429
00:26:20,600 --> 00:26:22,360
whole surface and the critical
parts.

430
00:26:22,360 --> 00:26:25,800
And then also points at the
closer to the surface and the

431
00:26:25,800 --> 00:26:27,720
volume are represented
correctly.

432
00:26:27,800 --> 00:26:30,760
All this, this type of stuff.
But there is a certain amount of

433
00:26:30,760 --> 00:26:33,160
points where you really capture
the whole phenomenon.

434
00:26:33,680 --> 00:26:37,240
And that's the, the, the, the
resolution invariant.

435
00:26:37,240 --> 00:26:39,640
So to say.
Obviously, if you go lower with

436
00:26:39,640 --> 00:26:42,920
the point to sample, you lose a
bit of representation.

437
00:26:42,920 --> 00:26:46,640
So you don't, you don't resolve
the full physics.

438
00:26:46,640 --> 00:26:50,160
So you're never able to really
recover the whole physics.

439
00:26:50,160 --> 00:26:53,280
But there is a, a certain amount
of points you need.

440
00:26:53,280 --> 00:26:56,720
And this is for, for each type
of physics problem that you have

441
00:26:56,720 --> 00:26:58,800
the full information in the
network.

442
00:26:59,240 --> 00:27:01,640
And then obviously depends how
good your network is.

443
00:27:02,080 --> 00:27:05,640
And then obviously it makes much
more sense to not have a network

444
00:27:05,640 --> 00:27:09,880
which again spits out this 128
points, but which conceptually

445
00:27:09,880 --> 00:27:13,040
is able to spit out as many
points as you want because

446
00:27:13,040 --> 00:27:16,840
that's the new operator, right?
That you can really get every

447
00:27:16,840 --> 00:27:19,040
point on the output function if
if needed.

448
00:27:20,120 --> 00:27:23,280
And this is something which is
which is very, very interesting

449
00:27:23,280 --> 00:27:25,760
in CFD.
Yeah.

450
00:27:25,760 --> 00:27:29,880
So I guess, again, correct me if
I'm wrong, if I'm distilling

451
00:27:29,880 --> 00:27:33,480
this correctly, is you're saying
if the surface in reality has

452
00:27:33,520 --> 00:27:38,440
8,000,000 points or whatever the
exact number is, you're saying

453
00:27:38,440 --> 00:27:43,480
that you should be able by
training and sampling

454
00:27:43,480 --> 00:27:48,880
progressively different points,
you shouldn't need more than

455
00:27:48,880 --> 00:27:54,200
256,000 to actually do as good
job as 8 million.

456
00:27:54,520 --> 00:27:58,240
Whereas I guess we've the
analogy.

457
00:27:58,480 --> 00:28:01,520
I suppose what I'm trying to
pick out is with mesh graph Nets

458
00:28:01,520 --> 00:28:05,720
or with graph neural Nets,
people could in theory take a

459
00:28:05,720 --> 00:28:08,920
million, but they would normally
like down sample, but they're

460
00:28:08,920 --> 00:28:14,400
down sampling at least.
Correct me if I'm wrong, you are

461
00:28:14,400 --> 00:28:17,960
then changing the problem
statement in this it it will

462
00:28:17,960 --> 00:28:22,120
give you a different answer and
and that I know was certain

463
00:28:22,120 --> 00:28:24,120
experience that we had and
others have had where you go

464
00:28:24,120 --> 00:28:28,440
well, how much do I downsample
and where do I pick the point to

465
00:28:28,440 --> 00:28:30,880
downsample?
And so a lot of people take an

466
00:28:30,880 --> 00:28:35,520
8,000,000 cell can't fit in AGP
memory and take 500,000 but

467
00:28:35,520 --> 00:28:40,320
they're worried when they do
inference they then would have

468
00:28:40,320 --> 00:28:44,040
to give the exact same mesh
distribution or if they give a

469
00:28:44,040 --> 00:28:47,080
different mesh distribute.
Do do you know what I'm getting?

470
00:28:47,080 --> 00:28:48,960
At that's lovely.
I mean you, you say the most

471
00:28:48,960 --> 00:28:51,760
important point, which I forgot.
Obviously if you have to

472
00:28:51,760 --> 00:28:54,880
calculate something with drag or
lift coefficient for which you

473
00:28:55,240 --> 00:28:57,320
you really need to get the
correct value.

474
00:28:57,320 --> 00:29:00,600
You need the full simulation
mesh, the full 8 to 9 million

475
00:29:00,960 --> 00:29:03,480
surface mesh.
And obviously the network has to

476
00:29:03,480 --> 00:29:07,040
give you this full mesh because
otherwise you're never able to

477
00:29:07,040 --> 00:29:09,440
calculate correct drag and lift
coefficient.

478
00:29:09,800 --> 00:29:12,920
But the network is giving you
this full network idea.

479
00:29:12,920 --> 00:29:17,000
Sorry, this full drag and lift
coefficient on the full surface

480
00:29:17,000 --> 00:29:23,440
mesh, no matter if you give it
64,128 thousand 256,000 input

481
00:29:23,440 --> 00:29:27,440
points, it's just at some in
number of input points the

482
00:29:27,440 --> 00:29:30,240
output is stagnating or the
performance is stagnating

483
00:29:30,240 --> 00:29:34,200
because you have reached the
maximum performance and the the

484
00:29:34,200 --> 00:29:38,440
input and and obviously you can
also give 8 million input

485
00:29:38,440 --> 00:29:40,720
points.
The the the self attention will

486
00:29:40,720 --> 00:29:43,720
be terribly slow and I don't
know what what parallelization

487
00:29:43,720 --> 00:29:47,120
tricks you need to do, but
nevertheless it will give you

488
00:29:47,120 --> 00:29:50,840
the same answer in the output.
So you can trustically reduce

489
00:29:50,840 --> 00:29:53,960
what you give as input as long
as the output still gives you

490
00:29:53,960 --> 00:29:56,280
the correct answer.
And that's, that's the magic we

491
00:29:56,360 --> 00:29:59,560
we got from transformer, which
we cannot do with with other

492
00:29:59,560 --> 00:30:01,760
methods for.
So for other methods we really

493
00:30:01,760 --> 00:30:04,800
had to make the problem much
easier, which is obviously

494
00:30:04,800 --> 00:30:12,520
giving you the wrong physics.
So if we down move on to and I

495
00:30:12,520 --> 00:30:16,360
guess this is the the novelty
that I saw in your approach and

496
00:30:16,400 --> 00:30:18,240
and I thought it was good to
double click on a bit.

497
00:30:18,240 --> 00:30:25,320
It's why the Transformer, what
specifically is that

498
00:30:25,320 --> 00:30:31,120
architecture giving you versus
alternatives because yeah, and

499
00:30:31,120 --> 00:30:33,640
how and how that relate, if you
could maybe go into a little bit

500
00:30:33,640 --> 00:30:36,640
because Transformers, I think
most people know conceptually,

501
00:30:36,640 --> 00:30:40,480
you know, tokens, etcetera,
LLMS, but maybe not in the

502
00:30:40,480 --> 00:30:44,960
context of CFD.
Yeah, so transformer are

503
00:30:45,400 --> 00:30:50,840
tremendously flexible,
tremendously optimised and and

504
00:30:50,840 --> 00:30:55,480
and tremendously well understood
powerhouse power work, power

505
00:30:55,480 --> 00:30:57,320
horse in, in, in deep learning,
right.

506
00:30:57,320 --> 00:31:00,200
They are they the the engine
behind large language model

507
00:31:00,200 --> 00:31:02,600
computer vision and and all
these tricks are already done

508
00:31:04,800 --> 00:31:08,920
and they have a few certain
properties which are super nice.

509
00:31:08,920 --> 00:31:12,760
The 1st is the the invariance
with respect to sequence length.

510
00:31:13,160 --> 00:31:17,280
So no matter how long the the
sequence length is which

511
00:31:17,280 --> 00:31:20,720
corresponds how many points you
input to the network, it will do

512
00:31:20,800 --> 00:31:23,720
the same calculations, right?
Which is very important for the

513
00:31:23,720 --> 00:31:31,840
scaling properties.
And then they have this how to

514
00:31:31,840 --> 00:31:35,920
say so I call it discretization
conversion.

515
00:31:35,920 --> 00:31:41,280
So if you sample more points in
a in an error in an area, it is

516
00:31:41,960 --> 00:31:45,320
basically it will converge to
some to some field.

517
00:31:46,080 --> 00:31:48,440
So, so the more points to
sample, the, the better the

518
00:31:48,440 --> 00:31:52,240
resolution gets, which is
obviously extremely valid,

519
00:31:52,400 --> 00:31:55,800
important property for this new
operator paradigm.

520
00:31:57,400 --> 00:32:00,880
They they are made.
So the transformer paradigm is

521
00:32:00,880 --> 00:32:03,640
made and, and stress tested
again and again and again for

522
00:32:03,640 --> 00:32:06,800
scaling, meaning that you can
build larger models, that you

523
00:32:06,800 --> 00:32:11,440
can ingest larger data sets.
Yeah.

524
00:32:11,440 --> 00:32:14,280
And, and, and all these things
make them an ideal fit for us.

525
00:32:15,320 --> 00:32:17,880
I mean, I'm very happy that not
everyone is using Transformers.

526
00:32:17,880 --> 00:32:20,120
That gives us some edge, but
yeah.

527
00:32:21,200 --> 00:32:25,840
But specifically, So how do you
tokenize the problem then?

528
00:32:25,840 --> 00:32:29,160
How could people conceptually
understand, you know, mesh graph

529
00:32:29,160 --> 00:32:32,960
Nets conceptually was I have a
node which corresponds to my

530
00:32:32,960 --> 00:32:36,760
mesh, I have some edges and then
I take that like you're not

531
00:32:36,760 --> 00:32:41,400
taking a token per node clearly
or that wouldn't scale, right.

532
00:32:41,440 --> 00:32:47,000
So how could people conceptually
think of that sort of tokenizing

533
00:32:47,000 --> 00:32:51,360
the the surface or the volume?
I think that the way you have to

534
00:32:51,360 --> 00:32:54,760
approach this is not from a
physicist perspective, but from

535
00:32:54,760 --> 00:32:59,480
a computer vision perspective.
So what do people do in computer

536
00:32:59,480 --> 00:33:01,120
vision?
So I think when do, when you do

537
00:33:01,320 --> 00:33:05,800
something like image generation
where you type in a prompt, I

538
00:33:05,800 --> 00:33:09,960
want a horse in the, in the in
the woods with whatever.

539
00:33:10,720 --> 00:33:14,000
So you have some text and then
you mix that with some image

540
00:33:14,000 --> 00:33:17,000
which is generated, right?
So you have basically 2 streams

541
00:33:17,680 --> 00:33:22,040
and, and the second important
part is this concept of

542
00:33:22,040 --> 00:33:27,040
patching.
So the transformer for for text,

543
00:33:27,040 --> 00:33:29,800
it's very clear each word gets a
token.

544
00:33:30,360 --> 00:33:36,600
So each word in a sentence or
probably each each sign or

545
00:33:36,600 --> 00:33:41,240
whatever is tokenized for for
images, small patches of the

546
00:33:41,240 --> 00:33:43,800
image will correspond to a token
and a word.

547
00:33:44,720 --> 00:33:47,400
So everything is a token which
then can interact with each

548
00:33:47,400 --> 00:33:50,480
other.
But in in language, it's the

549
00:33:50,480 --> 00:33:55,080
token's words in, in envision,
the tokens is small patches of

550
00:33:55,080 --> 00:33:58,320
of an image.
So if we do the same thing in in

551
00:33:58,320 --> 00:34:03,200
CFD, what we actually have, we
have basically three type of of

552
00:34:03,200 --> 00:34:05,320
information.
We have the information of the

553
00:34:05,320 --> 00:34:07,840
geometry.
So in general, what geometry we

554
00:34:07,840 --> 00:34:11,960
have, we have the information of
the volume and we have the

555
00:34:11,960 --> 00:34:18,159
information of basically the
physics on the on the geometry.

556
00:34:18,440 --> 00:34:21,880
So it makes sense to treat this
as either two or three

557
00:34:22,280 --> 00:34:26,480
modalities as we call it, right?
So as as you have text and and

558
00:34:26,480 --> 00:34:32,440
and image and then you can
tokenize either small areas,

559
00:34:32,880 --> 00:34:35,520
which makes sense.
If you if you look, try to

560
00:34:35,520 --> 00:34:39,760
encode the geometry so small
neighboring areas get pulled

561
00:34:39,760 --> 00:34:43,159
into one token so that you have
this information which gets

562
00:34:43,199 --> 00:34:47,719
locally aggregated.
Or you can for, for actually

563
00:34:47,719 --> 00:34:51,760
doing physics, you just take
some samples of the mesh and see

564
00:34:51,760 --> 00:34:54,199
them as a token.
And then you have the same

565
00:34:54,199 --> 00:34:56,840
setup, you have different
tokens, they represent different

566
00:34:56,840 --> 00:34:59,800
physic objects or physic parts
of the physics.

567
00:35:00,080 --> 00:35:03,200
And then similar to what people
do in computer vision, you let

568
00:35:03,200 --> 00:35:04,600
those tokens speak to each
other.

569
00:35:04,600 --> 00:35:09,640
You like, like this volume
tokens between each other and

570
00:35:09,640 --> 00:35:11,920
the volume tokens to the surface
and so on and so forth.

571
00:35:12,520 --> 00:35:16,600
So it's really borrowing a lot
of concepts from computer vision

572
00:35:16,600 --> 00:35:23,160
because that's what it works.
So if I understand it correctly,

573
00:35:23,520 --> 00:35:30,280
essentially what you're doing is
patching.

574
00:35:30,280 --> 00:35:34,360
So if you were to look at the
surface mesh and if you have,

575
00:35:34,760 --> 00:35:37,160
you know, visually, if you
looked at it and you had a clump

576
00:35:37,160 --> 00:35:44,680
of 5 by 5, then you, you would
have, you know, 25 cells and

577
00:35:44,680 --> 00:35:48,000
however many edges and nodes,
you would say, OK, well, that

578
00:35:48,280 --> 00:35:50,960
can just become one token and
then you go into the next, the

579
00:35:50,960 --> 00:35:53,160
next.
So if you, so essentially you're

580
00:35:53,160 --> 00:35:58,880
able to go from 8,000,000 points
and instead of taking 8 million

581
00:35:58,880 --> 00:36:02,480
tokens, you would have you know,
10,000 or what whatever the

582
00:36:02,480 --> 00:36:06,640
number is, is that.
And the assumption is that the

583
00:36:06,640 --> 00:36:13,760
differences within that patch
should be small enough to be not

584
00:36:13,760 --> 00:36:16,440
important.
If you were to take too big a

585
00:36:16,440 --> 00:36:20,760
patch, I assume you would lose
some accuracy because you're

586
00:36:20,760 --> 00:36:23,760
losing some local information.
Would that be fair to say there

587
00:36:23,760 --> 00:36:25,840
comes a cut off point?
Yeah.

588
00:36:25,840 --> 00:36:28,560
So if you if you talk about
representing the geometry, what

589
00:36:28,560 --> 00:36:30,600
you correctly said, you can do
two things, right.

590
00:36:30,600 --> 00:36:33,120
You can treat each point of the
geometry individually.

591
00:36:33,440 --> 00:36:38,760
It turns out this, this drive ML
data set and so on so forth.

592
00:36:38,760 --> 00:36:42,120
The information is so rich.
So in order to really represent

593
00:36:42,120 --> 00:36:45,200
the geometry, you need roughly,
I don't know, half a million

594
00:36:45,200 --> 00:36:47,560
points.
So it makes sense to pool them

595
00:36:47,560 --> 00:36:51,360
before, to aggregate them
locally, and then to use those

596
00:36:51,360 --> 00:36:54,600
as, as, as tokens.
That's just computationally much

597
00:36:54,600 --> 00:36:58,400
more efficient to have some sort
of pooling that the number of

598
00:36:58,400 --> 00:37:01,520
tokens is not exploding.
You could also obviously use

599
00:37:02,040 --> 00:37:07,920
representatives of the geometry
as your points, but it makes

600
00:37:07,920 --> 00:37:11,880
sense to pull them before.
And the same in the volume,

601
00:37:11,880 --> 00:37:13,840
right?
You're essentially, instead of

602
00:37:13,840 --> 00:37:20,760
taking 130 million nodes, you're
breaking them into patches which

603
00:37:20,760 --> 00:37:24,560
themselves take into.
Yeah, the The funny thing is so

604
00:37:24,560 --> 00:37:27,440
so there is this tool.
So we always have the the number

605
00:37:27,440 --> 00:37:30,360
of points which it takes to
represent physics correctly.

606
00:37:31,920 --> 00:37:34,000
This is 2 slightly different
things.

607
00:37:34,000 --> 00:37:37,000
So for the geometry we really
need a lot of tokens in order to

608
00:37:37,000 --> 00:37:39,600
really capture the geometry
because the geometry is

609
00:37:39,600 --> 00:37:41,280
influencing the physics and so
on and so forth.

610
00:37:41,680 --> 00:37:44,200
For the volume we don't need so
many points.

611
00:37:44,200 --> 00:37:48,920
So we we only need roughly
128,000 or or even less to

612
00:37:48,920 --> 00:37:52,840
really capture the physics.
So that's we are totally fine to

613
00:37:53,120 --> 00:37:56,080
pick those points individually
and actually performance is not

614
00:37:56,080 --> 00:38:00,880
degrading much if we go to 32 or
64,000 that will get more

615
00:38:00,880 --> 00:38:03,080
interesting if we get even
larger data sets.

616
00:38:03,560 --> 00:38:05,040
But then you can also use some
tricks.

617
00:38:05,040 --> 00:38:08,400
So there it's obviously fully
fine to use individual points.

618
00:38:09,920 --> 00:38:19,440
So in that context, is it fair?
A mental model is that you're

619
00:38:19,440 --> 00:38:23,440
saying that the number of points
that are needed for a numerical

620
00:38:23,440 --> 00:38:27,800
solver should not be a one to
one mapping.

621
00:38:28,720 --> 00:38:32,560
That you shouldn't think just
because I need half a billion

622
00:38:32,560 --> 00:38:36,240
points to solve my PDE, that I
should.

623
00:38:36,680 --> 00:38:40,480
I should need half a billion
points for my mission, my

624
00:38:40,480 --> 00:38:43,240
machine learning model to learn
the behaviour that we need to

625
00:38:43,240 --> 00:38:49,240
break from this concept of
almost one to one mapping.

626
00:38:49,240 --> 00:38:53,680
It's that part of your argument,
because that's not most people

627
00:38:53,680 --> 00:38:57,400
that I think assumed in the GNN
that I need to have this sort of

628
00:38:57,400 --> 00:39:01,640
1 to one mapping that I'm, I'm
sort of taking my notes and I'm

629
00:39:01,640 --> 00:39:04,840
putting that into my ML and
that's the most logical way of

630
00:39:04,840 --> 00:39:07,360
learning the problem.
Yeah.

631
00:39:07,360 --> 00:39:10,040
So I, I really like those
questions, I have to say.

632
00:39:10,040 --> 00:39:13,680
So I would my argument is, and I
mean you tell me you're the CFT

633
00:39:13,680 --> 00:39:19,040
expert, but I would say what a
model has to be able to do is it

634
00:39:19,040 --> 00:39:20,800
has to produce the simulation
mesh.

635
00:39:20,800 --> 00:39:23,840
So it has to give you
predictions on the simulation

636
00:39:24,000 --> 00:39:27,080
mesh both on the on the surface
and in the volume.

637
00:39:27,400 --> 00:39:30,440
And this is not only cars, it
should be hold for airplanes and

638
00:39:30,440 --> 00:39:33,120
so on and so forth.
And if we, if we think of meshes

639
00:39:33,120 --> 00:39:37,720
which approach a billion of mesh
points and where this is #1

640
00:39:38,600 --> 00:39:42,080
condition, we cannot, we cannot
use paralisation here because

641
00:39:42,080 --> 00:39:45,480
then we end up using thousand
GPU's and this is not scaling

642
00:39:45,480 --> 00:39:48,080
very well.
I'm looking at the actual

643
00:39:48,080 --> 00:39:52,200
information content.
So it has to be the model has to

644
00:39:52,200 --> 00:39:56,560
be able to produce this this
rich outputs in order to be a

645
00:39:56,880 --> 00:40:02,640
good tool to use.
But on the other hand you don't

646
00:40:02,640 --> 00:40:04,840
need this huge output for
training.

647
00:40:04,840 --> 00:40:07,720
You can train on much much
smaller sub sampled version of

648
00:40:07,720 --> 00:40:10,240
the same problem.
However, you have to make sure

649
00:40:10,360 --> 00:40:14,320
with null operator learning with
function approximation that

650
00:40:14,320 --> 00:40:17,720
those models you trained are
then in inference or in test

651
00:40:17,720 --> 00:40:20,160
time.
Able to give you the full answer

652
00:40:20,160 --> 00:40:22,960
on the full simulation measure.
I think this is the true CFD

653
00:40:22,960 --> 00:40:27,400
problem or the true learning
problem similar to the error 5

654
00:40:27,400 --> 00:40:30,960
setup in Feather modelling which
we have to do in CFD.

655
00:40:32,800 --> 00:40:37,720
So you yeah, so this is an
interesting point that I it's

656
00:40:37,720 --> 00:40:40,680
good to maybe briefly discuss
which is at imprint time.

657
00:40:41,000 --> 00:40:45,000
So we discussed about training
time, you know what you take,

658
00:40:45,000 --> 00:40:47,960
how many of the points, how many
do you agglomerate, you know,

659
00:40:47,960 --> 00:40:50,360
etcetera, etcetera.
But I guess what you're getting

660
00:40:50,360 --> 00:40:57,240
to now is the inference time,
which is something that maybe is

661
00:40:57,240 --> 00:41:00,040
not so obvious to some people,
but for me was always a bit of a

662
00:41:00,040 --> 00:41:06,200
conceptual challenge, which is
how do you generate the mesh to

663
00:41:06,200 --> 00:41:12,480
do the inference on.
So normally if you are running a

664
00:41:12,480 --> 00:41:15,480
normal CFD problem or you're
taking the driver ML data set,

665
00:41:15,480 --> 00:41:20,680
you split it into a train and
test, but both of them already

666
00:41:20,680 --> 00:41:23,560
have a mesh generated.
You know, we're giving it you.

667
00:41:24,200 --> 00:41:27,080
So you already have the mesh.
But the real life engineering

668
00:41:27,080 --> 00:41:30,720
problem is I want to predict a
new geometry that I've never

669
00:41:30,720 --> 00:41:33,600
seen before.
And the question is, do I have

670
00:41:33,600 --> 00:41:37,880
to mesh it like I would mesh
ACFD problem or can I just take

671
00:41:37,880 --> 00:41:42,240
the CAD or can I just do some
arbitrary case?

672
00:41:43,600 --> 00:41:47,120
Am I understanding it right that
this is sort of your argument of

673
00:41:47,120 --> 00:41:53,600
the neural operator or the field
sort of base approach That if I

674
00:41:54,080 --> 00:41:59,160
take a new geometry of a new car
that I haven't done anything

675
00:41:59,160 --> 00:42:07,400
before on and I mesh it 5
different ways and then ask your

676
00:42:07,400 --> 00:42:11,280
model to give me an A
prediction, it should

677
00:42:11,280 --> 00:42:16,200
essentially give pretty much the
same answer for those five

678
00:42:16,200 --> 00:42:21,240
different ways of meshing it.
Yes, I think this is the true.

679
00:42:22,320 --> 00:42:25,520
To an end, to a point.
I mean, two things.

680
00:42:26,160 --> 00:42:30,000
One, one thing I haven't said
so, so I I was talking about

681
00:42:30,520 --> 00:42:32,800
the, the difference in in
training and inference mesh.

682
00:42:33,200 --> 00:42:36,840
I also should say for for us
it's very important that the

683
00:42:36,840 --> 00:42:40,640
algorithm we are developing
works for for cat geometry.

684
00:42:40,640 --> 00:42:44,560
So if you input interference
only the cat geometry, you're

685
00:42:44,560 --> 00:42:49,520
able to obtain the full surface
and full volume without giving

686
00:42:49,560 --> 00:42:53,160
any information of the surface
and volume mesh in inference.

687
00:42:53,320 --> 00:42:56,360
The only thing the model gets
the CFD input.

688
00:42:57,320 --> 00:43:01,240
Obviously the performance
slightly degrades, but I'm

689
00:43:01,240 --> 00:43:05,560
amazed how well that works.
And you can do that by using,

690
00:43:05,560 --> 00:43:08,160
obviously in training that the
simulation mesh and, and, and

691
00:43:08,160 --> 00:43:11,440
using some tricks on the
simulation mesh that because

692
00:43:11,440 --> 00:43:13,480
obviously you need this
information in some way in

693
00:43:13,480 --> 00:43:16,720
training.
But that you can tell your model

694
00:43:17,360 --> 00:43:21,960
how to work with, with basically
structured points in the in the

695
00:43:21,960 --> 00:43:25,280
volume and, and certain points
on the on the surface such that

696
00:43:25,280 --> 00:43:28,240
interference it's the model is
able to deal without any of

697
00:43:28,240 --> 00:43:31,640
those meshings.
And I think this is somehow the

698
00:43:31,640 --> 00:43:35,280
true power of, of these deep
learning models because they

699
00:43:35,280 --> 00:43:39,160
cannot only scale.
You can, you can think of that

700
00:43:39,160 --> 00:43:41,560
the whole CFD compressed
simulation is suddenly

701
00:43:41,560 --> 00:43:44,080
compressed to, to neural network
weights.

702
00:43:44,160 --> 00:43:48,840
You just have to ask them the,
the, the model, which region in

703
00:43:48,840 --> 00:43:52,000
the, in the 3D you want to have
the prediction.

704
00:43:52,080 --> 00:43:56,280
But you can also do that from
pure COD inputs.

705
00:43:56,720 --> 00:44:00,560
And so, so no storing, no
whatever, right?

706
00:44:01,040 --> 00:44:03,600
And yeah, this is truly
exciting, I would say.

707
00:44:04,920 --> 00:44:08,440
Yeah, I think that's one of the
bits that I think is the true

708
00:44:08,440 --> 00:44:10,520
test.
Because if you have to generate

709
00:44:10,520 --> 00:44:16,080
a mesh, the real time element of
the prediction starts to go.

710
00:44:16,080 --> 00:44:19,240
Because for many people,
generating the mesh, the volume

711
00:44:19,240 --> 00:44:21,640
of the surface is itself quite a
challenge.

712
00:44:22,680 --> 00:44:26,720
And I think most people would
prefer to go straight from a CAD

713
00:44:26,720 --> 00:44:34,160
geometry or, or at least be less
sensitive, you know, to, to, to,

714
00:44:34,200 --> 00:44:36,280
to the mesh.
The, the only conceptual

715
00:44:36,280 --> 00:44:38,000
challenge to this, and I, I
don't know if we've even

716
00:44:38,000 --> 00:44:41,760
discussed this before, but I'll
just throw it out there, which I

717
00:44:41,760 --> 00:44:46,320
think still is part of the
challenge in people's heads is

718
00:44:46,840 --> 00:44:50,800
if you do ACFD simulation, you
would expect there to be a

719
00:44:50,800 --> 00:44:54,640
difference with the mesh, right?
It it should give a difference

720
00:44:55,040 --> 00:44:59,240
if I run that driver ML with a a
grid that's half as course or

721
00:44:59,240 --> 00:45:02,200
twice as fine.
I want it to give a difference

722
00:45:02,520 --> 00:45:05,960
because by having a course of
mesh I have higher numerical

723
00:45:05,960 --> 00:45:08,520
error which should affect the
flow field.

724
00:45:08,560 --> 00:45:15,120
And if I have a much finer mesh
I should see a difference

725
00:45:15,120 --> 00:45:17,040
because I'm reducing the
numerical error.

726
00:45:17,560 --> 00:45:19,840
I think where it gets a little
bit harder with the notion of

727
00:45:19,840 --> 00:45:23,920
sort of resolution independence,
is it sort of breaks from that

728
00:45:23,920 --> 00:45:27,280
mindset that, that I should be
able to give you a mesh that's

729
00:45:27,280 --> 00:45:29,040
twice as fine or twice as
coarse.

730
00:45:29,040 --> 00:45:32,320
And yet I get the same answer.
It it, it sort of messes with

731
00:45:32,320 --> 00:45:35,040
the CFD mind because you think
it should be different because I

732
00:45:35,560 --> 00:45:39,760
I'm expecting in my normal PD
solver for it to be different

733
00:45:40,240 --> 00:45:43,440
but it's not remember.
Oh, it's, yeah.

734
00:45:43,440 --> 00:45:45,920
I mean, actually this was one of
the reasons how we developed

735
00:45:45,920 --> 00:45:50,440
this, this anchor tokens, we
call them anchor tokens approach

736
00:45:50,440 --> 00:45:54,040
because this very much is, I
think the neural analogy of what

737
00:45:54,040 --> 00:45:56,600
you're saying.
So what these anchor tokens are

738
00:45:56,600 --> 00:45:58,800
and basically they are the
reasons how we can scale.

739
00:45:59,000 --> 00:46:03,640
So we train on a selected set of
tokens on the surface and on the

740
00:46:03,640 --> 00:46:06,600
volume.
And those tokens, as we already

741
00:46:06,600 --> 00:46:08,200
said, they need to capture the
physics.

742
00:46:08,600 --> 00:46:12,600
So it's, it needs to be a
certain amount, but they are

743
00:46:12,600 --> 00:46:15,080
sufficiently well to train with
self attention everything on

744
00:46:15,080 --> 00:46:16,840
Transformers and in fear and
cyber.

745
00:46:17,520 --> 00:46:19,520
What we do, we, we, we pick
these tokens.

746
00:46:19,520 --> 00:46:23,240
Those are representative points
in the whole 3D volume from

747
00:46:23,240 --> 00:46:26,320
surface and the volume.
And then basically those tokens

748
00:46:26,320 --> 00:46:28,440
span the weights of the keys and
values.

749
00:46:28,440 --> 00:46:31,160
So in the in the Transformer,
you have three different blocks

750
00:46:31,160 --> 00:46:35,640
and those span two blocks.
So you do a first pass of the

751
00:46:35,640 --> 00:46:38,320
model with those tokens and then
this is fixed.

752
00:46:38,320 --> 00:46:41,440
So you basically have what
you've built is you've built a

753
00:46:41,440 --> 00:46:46,440
model which which is is a
representation of of your full

754
00:46:46,440 --> 00:46:51,280
flow and and then you can ask
for specific points in the

755
00:46:51,280 --> 00:46:55,080
volume or on the surface to get
another value.

756
00:46:55,080 --> 00:46:57,800
And these points can be
arbitrary points, any point.

757
00:46:58,080 --> 00:47:01,920
And this point will just run
through the the model with all

758
00:47:01,920 --> 00:47:06,080
these weights already fixed.
So, so it only like a small part

759
00:47:06,080 --> 00:47:09,080
of the model is adjusted and it
will give you an output.

760
00:47:09,720 --> 00:47:13,680
And obviously the more anchor
points you have, the finer this

761
00:47:14,120 --> 00:47:17,000
this resolution gets, the better
this prediction are obviously

762
00:47:17,000 --> 00:47:18,360
until a certain threshold,
right.

763
00:47:18,840 --> 00:47:22,480
So if you think of 16,000 points
in the volume, you can still

764
00:47:22,480 --> 00:47:26,520
reconstruct the full flow field,
but it's it's not the full

765
00:47:26,520 --> 00:47:30,600
accuracy.
If you go to 32,064 thousand at

766
00:47:30,600 --> 00:47:33,560
some point, if the training
allowed it, you will have the

767
00:47:33,560 --> 00:47:36,800
full field covered.
And then with these, with these

768
00:47:36,800 --> 00:47:41,160
points, you can get arbitrary
find resolution and similar to

769
00:47:41,160 --> 00:47:44,480
in CFD and you and you cover all
the the points in space.

770
00:47:44,760 --> 00:47:49,400
I think I see this anchor points
really as as some, some some

771
00:47:49,400 --> 00:47:52,480
sort of of of this computation
cells in CFD.

772
00:47:52,840 --> 00:47:56,040
Obviously the the discrepancy
between number of mesh cells you

773
00:47:56,040 --> 00:47:59,560
need in CFD to get the
simulation to converge towards

774
00:47:59,640 --> 00:48:02,120
anchor points you need in order
to cover the full mesh.

775
00:48:02,120 --> 00:48:05,480
Is is trusted different?
But it's just because AI is

776
00:48:05,480 --> 00:48:07,800
trusted different, different
than the Mac simulation.

777
00:48:09,000 --> 00:48:11,680
But there is this analogy.
And obviously if you if you

778
00:48:11,680 --> 00:48:14,280
don't use any of these points,
it's very, very hard to

779
00:48:14,280 --> 00:48:17,000
construct these dynamics.
Yeah.

780
00:48:17,640 --> 00:48:20,720
I think conceptually some of the
interesting challenge of that

781
00:48:20,720 --> 00:48:29,000
is, correct me if I'm wrong, but
you, your ground truth is your

782
00:48:29,000 --> 00:48:32,400
train data.
So you feel that if you match

783
00:48:32,400 --> 00:48:35,440
the train data, that's the best
it can be.

784
00:48:36,360 --> 00:48:40,960
Right.
But the training data itself is

785
00:48:40,960 --> 00:48:46,760
dependent on the mesh.
And so one of the, I guess

786
00:48:46,760 --> 00:48:52,760
research topics that people have
discussed is, well, how about

787
00:48:52,760 --> 00:48:55,920
multi fidelity?
So you, you know, you have some

788
00:48:55,920 --> 00:49:02,560
high fidelity data, some low
fidelity data, but how?

789
00:49:05,640 --> 00:49:08,360
I'm so going off a tangent on
this a little bit, But one of

790
00:49:08,360 --> 00:49:12,200
the things I always say to
people is if you have a concept

791
00:49:12,200 --> 00:49:14,520
of like a foundational model
where you're saying, OK, I'm

792
00:49:14,520 --> 00:49:17,560
going to have a model that
predicts cars.

793
00:49:18,560 --> 00:49:24,200
I'm more and unless my mind is
just too static it it's a

794
00:49:24,200 --> 00:49:29,600
foundational model, but just for
the mesh design that you've

795
00:49:29,600 --> 00:49:32,360
picked and the CFD method you've
picked.

796
00:49:32,720 --> 00:49:36,280
If I then go and change that CFD
method and regenerate the driver

797
00:49:36,280 --> 00:49:39,760
ML with like a RANS data set,
I'm going to get a different

798
00:49:39,760 --> 00:49:43,880
answer.
So it's sort of hard to

799
00:49:43,880 --> 00:49:49,400
conceptually imagine the ML
being like a foundation model

800
00:49:49,400 --> 00:49:58,000
because the, the input data
itself is so dependent on the

801
00:49:58,000 --> 00:50:02,880
settings you have of the CFD.
So ideally it should really be

802
00:50:02,880 --> 00:50:06,400
like experimental date should be
the, the, the real ground truth.

803
00:50:07,200 --> 00:50:09,880
But it's, it's sort of a little
bit of a philosophical thing

804
00:50:09,880 --> 00:50:14,080
where, yeah, the ML model is
only as good as the training

805
00:50:14,080 --> 00:50:17,880
data that it has.
And so if now you've sort of

806
00:50:17,880 --> 00:50:21,720
solved the ML problem, well, I'm
not saying you've solved it.

807
00:50:22,520 --> 00:50:26,880
It's almost the data is the most
important bit now.

808
00:50:28,560 --> 00:50:33,480
I couldn't agree more.
So, so I, I, I would say talking

809
00:50:33,480 --> 00:50:36,840
about foundation models for
engineering is very, very

810
00:50:36,840 --> 00:50:40,560
dangerous because you're kind of
data agnostic in a way.

811
00:50:40,560 --> 00:50:44,760
You in the end, what goes into
these models is a huge amount of

812
00:50:44,760 --> 00:50:46,360
knowledge in this data
generation.

813
00:50:46,360 --> 00:50:50,760
And as you said, data on
numerics is not the ground

814
00:50:50,760 --> 00:50:53,520
truth.
It's just one version of, of, of

815
00:50:53,520 --> 00:50:55,920
our reality, right.
And, and machine learning is not

816
00:50:55,920 --> 00:51:00,280
replacing whatever this, this
version of reality, it's giving

817
00:51:00,280 --> 00:51:03,120
a certain new access to this
reality.

818
00:51:03,480 --> 00:51:07,760
So I, I think the power of
machine learning comes when you,

819
00:51:08,080 --> 00:51:11,920
when you really pick certain
areas where it's tremendously

820
00:51:11,920 --> 00:51:14,880
important to have this
surrogates and to have this vast

821
00:51:14,880 --> 00:51:19,480
iteration to have whatever and
then really think what it takes

822
00:51:19,480 --> 00:51:22,520
to build those surrogates.
But I think from a modelling

823
00:51:22,520 --> 00:51:27,680
perspective we are especially in
this static systems, we are

824
00:51:27,680 --> 00:51:31,920
pretty far that we can give them
enough data and high fidelity

825
00:51:31,920 --> 00:51:34,920
data and so on and so forth.
We can build these systems with

826
00:51:35,280 --> 00:51:38,600
reasonable errors compared to
the data is trained on.

827
00:51:39,560 --> 00:51:43,760
The question is then always how
to interact with numeric.

828
00:51:43,760 --> 00:51:48,320
So if depends on what you want,
but I think that the power and

829
00:51:48,320 --> 00:51:52,120
the true transformation comes
when when numerics and CFD

830
00:51:52,120 --> 00:51:55,720
really go and deep networks
really go hand in hand and not

831
00:51:55,720 --> 00:51:59,160
as as two different parties.
Maybe I should also say to that

832
00:51:59,160 --> 00:52:03,560
we were doing this drive ML data
set now for I don't know 10

833
00:52:03,560 --> 00:52:06,240
months like quite intensively
trying different things.

834
00:52:06,760 --> 00:52:09,840
We've reformulated the learning
problem ourselves a couple of

835
00:52:09,840 --> 00:52:11,800
times.
So what we actually want to

836
00:52:11,800 --> 00:52:14,040
train, what we actually want to
test and so on and so forth,

837
00:52:14,040 --> 00:52:17,920
because it's really, really hard
to because it's, it's just not a

838
00:52:17,920 --> 00:52:20,080
replacement of the CFD.
Your model is just doing

839
00:52:20,080 --> 00:52:23,160
something different and you have
to frame that there's a problem.

840
00:52:23,760 --> 00:52:27,400
And and so that's why I think
these two fields need to come

841
00:52:27,400 --> 00:52:30,920
closer together that these
questions are answered.

842
00:52:31,760 --> 00:52:35,160
Yeah, no, I and I was just
looking off to the side because

843
00:52:35,160 --> 00:52:39,560
I was just double checking if I
was correct that the error five

844
00:52:39,560 --> 00:52:45,280
data set is a combination of
actual satellite data, right

845
00:52:45,280 --> 00:52:49,080
observations with some sort of
modelling.

846
00:52:49,080 --> 00:52:53,280
So it is quite different in and
I think that's kind of my point

847
00:52:53,280 --> 00:52:55,760
that the weather stuff has
literally got what the

848
00:52:55,760 --> 00:52:59,120
satellites is seeing.
Therefore you could argue is the

849
00:52:59,120 --> 00:53:01,720
real thing because it is the
Earth.

850
00:53:01,720 --> 00:53:06,280
They saw what the weather's
doing on the Earth, whereas in

851
00:53:06,280 --> 00:53:11,680
the CFD it's, it is a, a just a
numerical simulation.

852
00:53:11,760 --> 00:53:16,720
And, and so that's the bit that
I still think is a bit of the

853
00:53:16,720 --> 00:53:23,680
limitation, but we are so sparse
when it comes to data.

854
00:53:24,800 --> 00:53:28,280
This is a bit of a segment into
a next question to you, which

855
00:53:28,280 --> 00:53:36,200
has long been the argument of
the pins people, which is are we

856
00:53:36,200 --> 00:53:40,400
ever going to have a vast
quantity of data to train sort

857
00:53:40,400 --> 00:53:44,360
of data-driven models?
Is is that a reasonable

858
00:53:44,360 --> 00:53:48,880
assumption or is it better to
say, well, actually we have

859
00:53:48,880 --> 00:53:54,120
pretty well defined equations
and rules and models and is it

860
00:53:54,120 --> 00:53:58,960
not just about using ML to solve
them faster?

861
00:53:59,040 --> 00:54:03,520
Like do we need, if we have such
a a sparsity of data, do we not

862
00:54:03,520 --> 00:54:08,120
need to include physics into the
models?

863
00:54:08,560 --> 00:54:13,160
Where do you sort of lie on that
conundrum of just generating

864
00:54:13,160 --> 00:54:17,560
more data and let the models of
the data or do we need to

865
00:54:17,560 --> 00:54:21,120
incorporate more physics into
these models?

866
00:54:21,360 --> 00:54:24,040
I'm very pragmatic in this.
I mean this, This was the same

867
00:54:24,040 --> 00:54:26,000
in weather modelling.
Everyone was talking about

868
00:54:26,520 --> 00:54:28,400
physics informed, not physics
informed.

869
00:54:28,600 --> 00:54:32,920
But people only talk for, for,
for that amount of time until

870
00:54:32,920 --> 00:54:35,560
they realise how heavy the data
loading is and how hard of a

871
00:54:35,560 --> 00:54:36,920
machine learning problem that
is.

872
00:54:37,720 --> 00:54:41,080
Because you don't want to talk
about the physics in form neural

873
00:54:41,080 --> 00:54:44,120
network if you have to, to check
it between loading gigabytes of

874
00:54:44,120 --> 00:54:46,720
data for one data point on the
GPO and what to do.

875
00:54:46,720 --> 00:54:50,200
And in order to get this correct
and in order to get the decent

876
00:54:50,200 --> 00:54:52,320
learning signal and so on and so
forth.

877
00:54:53,000 --> 00:54:57,520
And I, I see it the same for
this engineering simulations.

878
00:54:57,520 --> 00:55:01,440
We first have to be able to
really run simulations.

879
00:55:01,560 --> 00:55:04,440
It's scale.
And I'm not talking about 10,000

880
00:55:04,440 --> 00:55:06,280
cells.
I'm not talking about shape net,

881
00:55:06,280 --> 00:55:09,640
I'm not talking about even
driving at the drive ML.

882
00:55:09,640 --> 00:55:13,720
I'm talking about airplane
simulation and I'm talking about

883
00:55:14,160 --> 00:55:17,800
transient simulation and that.
And we are still far away from

884
00:55:17,800 --> 00:55:20,680
doing that both from how we
generate, how we store, how we

885
00:55:20,680 --> 00:55:24,160
train data.
And I think only if we, if we

886
00:55:24,160 --> 00:55:27,040
are able to really have a
workflow there and know what it

887
00:55:27,040 --> 00:55:29,600
takes and know where it breaks
and, and, and where all these

888
00:55:29,600 --> 00:55:33,440
problems are, we can then start
thinking about which type of

889
00:55:33,440 --> 00:55:38,520
physics to include and and which
not because, I mean, we, we also

890
00:55:38,520 --> 00:55:45,040
showed that in our paper, it's
very easy to build a diversions

891
00:55:45,040 --> 00:55:49,080
free model for, for vorticity,
because you can bake in the, the

892
00:55:49,080 --> 00:55:51,280
diversions free constraint,
their construction a hard

893
00:55:51,280 --> 00:55:54,640
constraint.
We try the pin loss for the, for

894
00:55:54,640 --> 00:55:58,520
the mass conservation, which is
much harder to just take the

895
00:55:58,520 --> 00:56:01,040
great performance.
It in principle works.

896
00:56:01,040 --> 00:56:04,600
And you would need some tricks
and I would say getting this to

897
00:56:04,600 --> 00:56:10,360
work is not the amount of time
which needs to do all the

898
00:56:10,360 --> 00:56:13,520
scaling.
But I would say that first the

899
00:56:13,520 --> 00:56:18,560
problems we have to solve first
are the really how we interact

900
00:56:18,560 --> 00:56:21,320
with numerics, how we scale to
these problems and then how we

901
00:56:21,320 --> 00:56:25,200
actually get industry ready.
Well, you, you mentioned the

902
00:56:25,200 --> 00:56:29,800
beginning, the sort of T to T +
1 analogy when it came to

903
00:56:29,800 --> 00:56:32,160
weather.
And clearly the bit that we

904
00:56:32,840 --> 00:56:35,080
missed out maybe in our
introduction is that the

905
00:56:35,080 --> 00:56:36,760
engineering, we're not doing
that right.

906
00:56:37,560 --> 00:56:41,840
You're going from yeah, straight
to a time average or something

907
00:56:42,040 --> 00:56:46,760
at least, at least in, at least
in, well, in in the sort of way

908
00:56:46,760 --> 00:56:50,280
that most of these data sets for
cars and planes have been to

909
00:56:50,280 --> 00:56:51,440
date.
Given.

910
00:56:52,320 --> 00:56:56,120
How would you change the
learning problem or the ML

911
00:56:56,120 --> 00:57:01,200
architecture if now you had a a
full time history and would that

912
00:57:01,200 --> 00:57:03,600
help hindrance not make a
difference?

913
00:57:04,600 --> 00:57:08,080
I mean same argument as as we
have in in the spatial.

914
00:57:08,080 --> 00:57:10,840
If you if you look at the the
neural network which has a few

915
00:57:11,440 --> 00:57:16,160
megabytes of weights and which
compresses, I don't know the,

916
00:57:16,240 --> 00:57:22,200
the GB CFD simulation or 10 or
20, I don't know how many GB

917
00:57:22,200 --> 00:57:26,280
actually CFD simulation, you can
also play the same game over

918
00:57:26,280 --> 00:57:28,880
time, right?
So neural network is just a

919
00:57:28,880 --> 00:57:33,560
very, very good way of of
storing this, this information.

920
00:57:34,080 --> 00:57:38,960
And if you somehow manage to to
store the temporal part in, in

921
00:57:39,000 --> 00:57:42,160
some sort of modelling, which
allows you to give you the the

922
00:57:42,160 --> 00:57:44,960
temporal answer to a certain
problem, this would be

923
00:57:45,480 --> 00:57:48,280
tremendously huge gain.
Because suddenly you can really

924
00:57:48,320 --> 00:57:53,080
simulate stuff in, in, in, in
real time, but also have the the

925
00:57:53,080 --> 00:57:55,960
ability to observe what's going
on and, and, and to really

926
00:57:55,960 --> 00:57:57,560
understand the temporal
dynamics.

927
00:57:57,960 --> 00:58:02,840
Whereas the training is just
somehow pumping this simulation

928
00:58:02,840 --> 00:58:06,240
into the neural network weights,
ideally without storing them.

929
00:58:06,240 --> 00:58:07,760
But that's that's very, very
hard.

930
00:58:07,760 --> 00:58:09,760
And then probably the next to,
next to next to step.

931
00:58:11,800 --> 00:58:18,160
But yeah, I think first step is
to get to get the decent neural

932
00:58:18,160 --> 00:58:21,240
network architectures for
temporal modelling, which we are

933
00:58:21,240 --> 00:58:22,600
actually working on to be
honest.

934
00:58:23,200 --> 00:58:25,440
OK.
And, and would you have the

935
00:58:25,440 --> 00:58:30,320
sense to be something similar
where you know, you may want to

936
00:58:30,320 --> 00:58:34,680
have a million time steps, but
that you'd expect that there

937
00:58:34,680 --> 00:58:39,760
would be some way of, I don't
know, is there, is there a sort

938
00:58:39,760 --> 00:58:47,240
of equivalency in terms of AD
inference?

939
00:58:47,680 --> 00:58:52,920
Ideally you wouldn't need to go
1,000,000 steps, right?

940
00:58:53,760 --> 00:58:59,040
Do you have a sense of how much
this is more question for people

941
00:58:59,040 --> 00:59:01,880
who are generating data, You
know, do I need to generate a

942
00:59:01,880 --> 00:59:04,680
million times steps of data so
that you can learn the million?

943
00:59:04,920 --> 00:59:08,720
And do you expect that then you
would be able to predict to the

944
00:59:08,720 --> 00:59:11,400
millionth but skip out every
hundred?

945
00:59:11,560 --> 00:59:14,680
Because you, you don't need to,
you know, to do that?

946
00:59:16,360 --> 00:59:20,000
Do you think you could learn the
temporal problem without every

947
00:59:20,000 --> 00:59:23,080
single time step from the
traditional PDE?

948
00:59:23,080 --> 00:59:24,640
Do you know what I'm trying to
get a sense of?

949
00:59:24,640 --> 00:59:27,720
Like the how the temporal
challenge would link to the

950
00:59:27,720 --> 00:59:31,760
spatial challenge?
I think it's it's very much

951
00:59:31,760 --> 00:59:33,640
relatable.
So in in space we see this

952
00:59:33,640 --> 00:59:37,880
phenomenon that you don't need
all the points to to to capture

953
00:59:37,880 --> 00:59:42,000
the physics and then depend.
Then you have a huge advantage

954
00:59:42,000 --> 00:59:45,520
in learning because you can
always sub sample different sets

955
00:59:45,520 --> 00:59:47,840
of these points which give you a
huge data augmentation.

956
00:59:48,160 --> 00:59:49,760
And you have the same in a
temporal domain.

957
00:59:49,760 --> 00:59:53,840
You don't need each time step in
order to to present the full

958
00:59:53,840 --> 00:59:57,360
physics to the neural network.
But obviously the the fine

959
00:59:57,360 --> 01:00:00,720
grain, the more the more fine
graining you have in your in

960
01:00:00,720 --> 01:00:04,520
your solution, the better you
can do the data augmentation in

961
01:00:04,520 --> 01:00:07,840
the temporal dimension.
But obviously machine learning

962
01:00:07,840 --> 01:00:12,200
models can do much, much bigger
time steps than numerical

963
01:00:12,200 --> 01:00:15,120
models.
And yeah, this is something

964
01:00:15,120 --> 01:00:17,880
which which is a true advantage.
They just don't work the same

965
01:00:17,880 --> 01:00:21,160
and they don't have stability
issues but also no stability

966
01:00:21,160 --> 01:00:24,520
guarantees well.
I was going to say, isn't that

967
01:00:24,800 --> 01:00:29,280
one of the problems, the roll
out problem that, you know, I

968
01:00:29,280 --> 01:00:34,400
guess if you do inference
inference, inference imprints,

969
01:00:35,640 --> 01:00:38,560
you know, isn't, isn't these
sort of instabilities going to

970
01:00:38,560 --> 01:00:41,680
build up or like, you know, how
do you constrain it?

971
01:00:42,520 --> 01:00:44,760
Is this where some of the
physics needs to be added?

972
01:00:44,760 --> 01:00:49,720
Do you think to sort of if you
need to do 100,000 iterations,

973
01:00:50,560 --> 01:00:53,640
how much are you going to
guarantee that just a small

974
01:00:53,640 --> 01:00:57,200
difference in that first is not
going to, you know, bifurcate

975
01:00:57,200 --> 01:01:01,160
into different solutions?
Yeah, I mean, there, there is 2

976
01:01:01,160 --> 01:01:03,080
to answer.
So that one is obviously you.

977
01:01:03,560 --> 01:01:05,680
Luckily people do video
generation nowadays.

978
01:01:05,680 --> 01:01:08,160
So we get a lot and lot of
tricks of video generation,

979
01:01:08,160 --> 01:01:11,880
which actually really, really
start to make this thing stable.

980
01:01:12,280 --> 01:01:15,200
And one of these, these
definitely modeling paradigms is

981
01:01:15,200 --> 01:01:17,960
this generative model that you
model the distribution over time

982
01:01:17,960 --> 01:01:20,480
and that you kind of make sure
the distribution always stays

983
01:01:20,480 --> 01:01:23,720
the same.
And, and this boils down to a

984
01:01:23,720 --> 01:01:25,520
bit of, of physics
understanding, right?

985
01:01:25,520 --> 01:01:28,080
If you, if you make sure the
distribution is the same, if you

986
01:01:28,360 --> 01:01:33,120
make sure that that future and
past they considered that there

987
01:01:33,120 --> 01:01:35,760
is a tension across the right
axis and so on and so forth, we

988
01:01:35,760 --> 01:01:39,400
get stability.
It's always depends on how you

989
01:01:39,400 --> 01:01:41,520
define physics.
For me, it's the information the

990
01:01:41,520 --> 01:01:44,480
model needs and the way you,
you, you, you interact with

991
01:01:44,480 --> 01:01:48,680
future and past and so on.
I truly think that this

992
01:01:48,680 --> 01:01:51,560
generative modelling is there's
the breakthrough in this, this

993
01:01:51,560 --> 01:01:54,440
rollout stabilities.
I remember when I ran first into

994
01:01:54,440 --> 01:01:59,000
this rollout problems, it was I
think 4 years ago and, and, and

995
01:01:59,000 --> 01:02:02,040
basically I, I did this paper
together with Daniel Worrell and

996
01:02:02,040 --> 01:02:04,920
I texted him and saying, hey, I
don't know, we have to do

997
01:02:04,920 --> 01:02:06,880
something.
Our rollouts always exploding.

998
01:02:06,880 --> 01:02:09,000
And then I have no idea.
And, and he said, well, yeah,

999
01:02:09,000 --> 01:02:12,320
there has to be some, some
tricks in literature and then

1000
01:02:12,320 --> 01:02:14,920
something smart and, and, and
definitely people have thought

1001
01:02:14,920 --> 01:02:16,760
about it.
And I was like, yeah, I cannot

1002
01:02:16,760 --> 01:02:19,160
find anything.
There is like, why is nobody

1003
01:02:19,160 --> 01:02:23,040
having this problem?
And it's basically just a very

1004
01:02:23,040 --> 01:02:26,080
different problem we have in in
in this engineering task that

1005
01:02:26,080 --> 01:02:29,880
the the input itself each time
step is so huge from information

1006
01:02:29,880 --> 01:02:32,880
content that every
autoaggressive model, every

1007
01:02:33,280 --> 01:02:36,760
every autoaggressive trick you
you had before is just not

1008
01:02:36,760 --> 01:02:40,920
working because they work with
much smaller input vectors.

1009
01:02:42,200 --> 01:02:45,960
But but now that computer vision
is going to videos and that

1010
01:02:45,960 --> 01:02:50,000
people are aware of this problem
and that frequency spectrum need

1011
01:02:50,000 --> 01:02:51,600
to be conserved and so on and so
forth.

1012
01:02:51,600 --> 01:02:56,600
This guess really a lot of lot
of progress.

1013
01:02:56,600 --> 01:02:59,520
And I would say that that the
having the frequency spectrum

1014
01:02:59,520 --> 01:03:02,160
stable is is is what made many
of these things stable.

1015
01:03:03,520 --> 01:03:07,360
So it sounds like in a way there
is still a lot of potential to

1016
01:03:07,360 --> 01:03:13,280
learn from the other advances.
Yeah, in ML, you know, the fact

1017
01:03:13,440 --> 01:03:15,960
there's always an analogy
problem I guess is what you're

1018
01:03:15,960 --> 01:03:18,080
saying.
There's the stuff you can learn

1019
01:03:18,080 --> 01:03:21,600
in terms of video generation,
text generation, weather,

1020
01:03:21,600 --> 01:03:25,200
climate week.
You can borrow ideas from other

1021
01:03:25,200 --> 01:03:28,120
fields, but it suggests that you
therefore need to have people

1022
01:03:28,120 --> 01:03:34,520
with a a broader ML background.
You know, if you just have a

1023
01:03:34,520 --> 01:03:39,520
fluids background, you may
struggle to solve this problem

1024
01:03:39,520 --> 01:03:41,880
because you need to know what or
you need to have an

1025
01:03:41,880 --> 01:03:44,400
understanding and an interest in
looking at what's video

1026
01:03:44,520 --> 01:03:46,320
generation doing or what's this
doing.

1027
01:03:46,640 --> 01:03:48,360
Would that be fair to say?
This is why it's a

1028
01:03:48,360 --> 01:03:52,160
multidisciplinary problem.
Almost all you need a team that

1029
01:03:52,160 --> 01:03:54,880
is multidisciplinary.
I mean that that's happening in

1030
01:03:54,880 --> 01:03:57,240
biotech and in other areas as
well, right?

1031
01:03:58,240 --> 01:04:02,200
We are just very in agnostic to
that and language and vision

1032
01:04:02,240 --> 01:04:06,160
because you don't need a
linguist to build the LLM and

1033
01:04:06,160 --> 01:04:10,120
you don't need basically, I
don't know even what the

1034
01:04:10,160 --> 01:04:13,440
computer vision like.
What's that analogy there?

1035
01:04:13,800 --> 01:04:18,400
But for, for, for, for all this
alpha fold and so on so forth.

1036
01:04:18,400 --> 01:04:22,920
The surely had domain experts on
the team and surely this is

1037
01:04:22,920 --> 01:04:25,880
what, what's the exciting part
about engineering you, you have

1038
01:04:25,880 --> 01:04:28,680
problems which are so hard to
correct that you need the top

1039
01:04:28,680 --> 01:04:31,160
notch machine learning people
that the top notch domain

1040
01:04:31,240 --> 01:04:33,400
people, they need to speak the
same language.

1041
01:04:33,840 --> 01:04:37,200
And even if you can borrow a lot
of concepts, they need to be

1042
01:04:37,200 --> 01:04:39,640
heavily adjusted.
Because the problem with this

1043
01:04:39,640 --> 01:04:42,840
large mesh is with what's the
input, what's the output?

1044
01:04:42,840 --> 01:04:45,720
There's this huge ratio
difference and, and, and, and

1045
01:04:45,720 --> 01:04:49,160
turbulence and whatnot.
This is just a, a totally

1046
01:04:49,160 --> 01:04:52,360
different piece to, to crack.
So that's what's the exciting

1047
01:04:52,360 --> 01:04:55,600
part here.
I, I wanted to maybe pivot it a

1048
01:04:55,600 --> 01:05:01,480
little bit to the process you
went in creating a start up.

1049
01:05:01,920 --> 01:05:05,320
I, I think a lot of people would
be interested to know and I'm

1050
01:05:05,320 --> 01:05:09,400
sure other people, you know, are
in a pub or, you know, alcohol

1051
01:05:09,400 --> 01:05:11,200
doesn't need to be involved, of
course, but you know, there's

1052
01:05:11,200 --> 01:05:15,440
somewhere and they've got ideas
and, but not many people

1053
01:05:15,440 --> 01:05:19,760
actually go out secure funding
and, you know, create a start

1054
01:05:19,760 --> 01:05:22,960
up.
So obviously not giving away any

1055
01:05:22,960 --> 01:05:25,840
sensitive details, but what was
that experience like for you?

1056
01:05:25,840 --> 01:05:29,440
Like how did it begin?
How did you deal with venture

1057
01:05:29,440 --> 01:05:33,120
capital companies, how to like?
Yeah, I'd be interested if you

1058
01:05:33,120 --> 01:05:37,400
could share some of your
experience in founding a start

1059
01:05:37,400 --> 01:05:39,400
up, some of the lessons learnt
maybe.

1060
01:05:39,640 --> 01:05:41,680
Yeah, so I mean, I, I was really
lucky.

1061
01:05:41,680 --> 01:05:44,880
So, so first of all, a lot of
knowledge comes up from my

1062
01:05:44,880 --> 01:05:47,560
university.
I, I have to know, I have to say

1063
01:05:47,560 --> 01:05:50,680
that the people like Benedict
Alkin and and Tobiascon Lachner

1064
01:05:50,680 --> 01:05:53,640
who are really, really pushing
the efforts, they have been with

1065
01:05:53,640 --> 01:05:58,240
me from, from day one actually.
Also people like Stefan Berka

1066
01:05:58,240 --> 01:06:02,400
here in Linz University is a
well known simulation expert who

1067
01:06:02,400 --> 01:06:06,360
who knows all these tricks.
When I met him the first day,

1068
01:06:06,360 --> 01:06:09,040
basically after returning to
Austria, we were 5 minutes and

1069
01:06:09,040 --> 01:06:11,720
we were both like super excited
that, that we do things in this

1070
01:06:11,720 --> 01:06:16,240
area.
And yeah, we, there was a lot of

1071
01:06:16,240 --> 01:06:18,840
momentum there.
We were at a different company

1072
01:06:18,840 --> 01:06:22,720
and XCI where at some point it
was clear that the simulation

1073
01:06:22,720 --> 01:06:26,880
team, which was was really
working well, should branch off.

1074
01:06:27,720 --> 01:06:31,640
I was also lucky to to have met
people like Dennis Houston mix

1075
01:06:31,640 --> 01:06:35,400
Mickelson's who who did this
before graded startups, very

1076
01:06:35,400 --> 01:06:37,600
successful startups and actually
knew what to do.

1077
01:06:38,040 --> 01:06:41,480
They were the first to listen to
me when I told them, hey, this

1078
01:06:41,480 --> 01:06:43,600
engineering like scaling these
things up.

1079
01:06:43,600 --> 01:06:46,680
What we do to this billion of
mesh problems is actually what

1080
01:06:47,080 --> 01:06:50,560
what industry needs.
But just give us some time.

1081
01:06:50,560 --> 01:06:54,320
We will figure that out.
And so I think it was all very

1082
01:06:54,320 --> 01:06:56,800
natural.
It was not planned, but it's

1083
01:06:56,800 --> 01:07:00,920
obviously a huge excitement if
you can, if you can do what

1084
01:07:00,920 --> 01:07:03,600
you're really burning for in,
in, in also in a, in a

1085
01:07:03,600 --> 01:07:06,640
commercial set up where you you
really can make an impact with

1086
01:07:06,640 --> 01:07:10,280
what you're doing.
So yeah, I would say in that

1087
01:07:10,280 --> 01:07:13,200
sense I was very lucky.
And I would also say that

1088
01:07:13,240 --> 01:07:16,680
especially in this area, we, we
see a lot of momentum generated

1089
01:07:16,680 --> 01:07:21,520
and we're not the last start up
and there is hopefully coming

1090
01:07:21,520 --> 01:07:25,080
more it, it, it there is a
momentum especially in Europe

1091
01:07:25,080 --> 01:07:28,080
now in in this area.
And yeah, just very exciting.

1092
01:07:28,080 --> 01:07:30,120
I would do it exactly the same
way again.

1093
01:07:31,400 --> 01:07:36,480
And so how do you think you
could have done this any other

1094
01:07:36,480 --> 01:07:38,680
way?
Do you think you could have come

1095
01:07:38,680 --> 01:07:42,440
up with what you're coming up at
a company or at the university

1096
01:07:42,800 --> 01:07:49,680
or is it is this is the start up
really the only way of I asked

1097
01:07:49,680 --> 01:07:52,240
this question because I asked
the same thing to Max Welling

1098
01:07:52,640 --> 01:07:57,160
and he was sort of saying, well,
yeah, start-ups is the only

1099
01:07:57,160 --> 01:08:00,640
place where this can happen.
The university's either too

1100
01:08:00,640 --> 01:08:03,240
small and doesn't have enough
funding or companies too big and

1101
01:08:03,240 --> 01:08:05,080
too slow.
And the start up is this sort

1102
01:08:05,080 --> 01:08:08,440
of, but it feels like in Europe,
but often a bit risk averse.

1103
01:08:09,600 --> 01:08:12,320
We almost feel like start-ups is
for somebody else.

1104
01:08:12,520 --> 01:08:15,320
But I'd be interested to know
your thoughts on.

1105
01:08:15,520 --> 01:08:18,920
That I'm not so sure if what
we're pulling off.

1106
01:08:19,240 --> 01:08:22,640
We just need the, the, the best
people and, and, and, and some

1107
01:08:22,640 --> 01:08:24,920
compute.
One thing which is really,

1108
01:08:25,080 --> 01:08:27,760
really, really, really hard is
the problems.

1109
01:08:28,160 --> 01:08:32,920
So it's actually we are choosing
our customer mostly by the, the,

1110
01:08:32,960 --> 01:08:37,240
the, the scale and size of the
problems they have because

1111
01:08:37,399 --> 01:08:40,760
they're really interesting
problems come from industry and

1112
01:08:40,760 --> 01:08:45,720
the more challenging and those
problems are, the better they

1113
01:08:45,720 --> 01:08:47,840
are for us and, and for
developing us.

1114
01:08:47,840 --> 01:08:51,560
So obviously we now have have a
customer which have huge

1115
01:08:51,560 --> 01:08:54,319
temporal problems where you
have, we have different

1116
01:08:54,319 --> 01:08:58,040
modalities interacting and this
can bootstrap our temporal

1117
01:08:58,040 --> 01:09:02,880
modelling capabilities and and
and and and and and and, and,

1118
01:09:02,920 --> 01:09:06,920
and this is something which you
wouldn't get in the industry I

1119
01:09:06,920 --> 01:09:10,080
at university.
So I would say it's a mixture.

1120
01:09:10,080 --> 01:09:13,920
I would say for me, it's, it's,
it really helps to get hands on

1121
01:09:13,920 --> 01:09:19,960
to the, to the real problems
and, and on, on the other side,

1122
01:09:19,960 --> 01:09:23,359
it's, it's just helps me to get
into touch what industry really

1123
01:09:23,359 --> 01:09:25,800
needs and, and not sit in my
ivory tower.

1124
01:09:26,200 --> 01:09:29,240
I wouldn't, I wouldn't stand
doing things where I know people

1125
01:09:29,240 --> 01:09:34,120
wouldn't use them.
So this is a good, yeah, good

1126
01:09:34,120 --> 01:09:36,240
mix.
But I also have to say there is

1127
01:09:36,240 --> 01:09:39,760
luckily some people like Neil
Ashton who generate publicly

1128
01:09:39,760 --> 01:09:43,279
available data sets that that
people can really start looking

1129
01:09:43,279 --> 01:09:47,319
at that We just have to give the
community the urgency that that

1130
01:09:47,319 --> 01:09:50,720
they really look at the hard
data sets and not a shape, no

1131
01:09:50,720 --> 01:09:53,720
car.
Yeah, it, it seems like the

1132
01:09:55,240 --> 01:09:59,880
there is definitely a movement
in the industry towards this

1133
01:09:59,880 --> 01:10:02,240
machine learning problems.
You know, every time I go to a

1134
01:10:02,240 --> 01:10:04,920
conference, there's more and
more people doing it, but it

1135
01:10:04,920 --> 01:10:10,320
still feels like we haven't
reached that inflection point.

1136
01:10:10,320 --> 01:10:16,800
It still feels like you're
either a start up or you're, I

1137
01:10:16,800 --> 01:10:19,120
don't know.
It still feels like we're not

1138
01:10:19,120 --> 01:10:25,640
fully at that yeah, inflection
point in terms of mass adoption

1139
01:10:25,640 --> 01:10:27,200
of this.
I feel like we're getting

1140
01:10:27,200 --> 01:10:28,920
closer.
I'm sure you feel the same.

1141
01:10:28,920 --> 01:10:35,120
You know, lots of industries
interested, but maybe it is the

1142
01:10:35,120 --> 01:10:39,240
data that's the issue.
Maybe it's the Era 5 was such a

1143
01:10:39,240 --> 01:10:43,440
large open data set that data
wasn't the issue and therefore

1144
01:10:43,440 --> 01:10:44,880
people just got into the
modelling.

1145
01:10:45,240 --> 01:10:50,320
Whereas I feel like now there's
been quite a lot of advances in

1146
01:10:50,320 --> 01:10:52,720
let's say Rd. car external
aerodynamics.

1147
01:10:53,600 --> 01:10:56,480
But Rd. car external
aerodynamics, I don't know

1148
01:10:56,600 --> 01:11:00,040
percentage wise, but it's it's a
small percentage of the entire

1149
01:11:00,040 --> 01:11:05,520
CFD domain.
And yeah, yeah, it feels like we

1150
01:11:05,520 --> 01:11:09,840
still need to convince maybe the
whole CFD community to do more,

1151
01:11:09,840 --> 01:11:15,920
to really advance it and make
this a more systematic, I guess.

1152
01:11:15,920 --> 01:11:19,200
Would that be true to say that
you can only work on the

1153
01:11:19,200 --> 01:11:21,160
problems that you updated for?
I mean, it sounds like an

1154
01:11:21,160 --> 01:11:24,240
obvious statement to make, but I
assume if you had 10 times more

1155
01:11:24,240 --> 01:11:27,640
data, you could potentially go
and hire more people and work on

1156
01:11:27,640 --> 01:11:29,440
more problems.
But you're not going to go and

1157
01:11:29,440 --> 01:11:33,520
hire more people and work on
other problems if there isn't

1158
01:11:34,480 --> 01:11:36,840
the data or the willingness to
do this.

1159
01:11:37,200 --> 01:11:40,160
And, and most people also don't
know what to, to optimise for,

1160
01:11:40,160 --> 01:11:42,480
right.
So, so I think the, the, the

1161
01:11:42,480 --> 01:11:45,560
industry or the, the, the real
problems need to need to set

1162
01:11:45,560 --> 01:11:48,800
the, the, the, the, the, the
stage they need to set the, the

1163
01:11:48,800 --> 01:11:51,000
machine learning problem.
Then we can start to think of

1164
01:11:51,000 --> 01:11:55,320
how to solve it.
Yeah, I, I, but there is a

1165
01:11:55,320 --> 01:11:57,840
momentum.
People are still reluctant

1166
01:11:57,840 --> 01:12:00,760
because getting a nervous paper
easier on a smaller data set

1167
01:12:00,760 --> 01:12:05,440
than on a larger data set.
But but I think there is quite

1168
01:12:05,440 --> 01:12:08,840
some movement and and and and
things are are really changing.

1169
01:12:09,440 --> 01:12:13,880
And I also from for me, it's,
it's kind of our philosophy that

1170
01:12:13,880 --> 01:12:16,640
to be a bit open source to show
our models to show what we're

1171
01:12:16,640 --> 01:12:21,600
doing to publish, because I
think this is this is the way to

1172
01:12:21,600 --> 01:12:26,440
go forward.
Closed, closed doors policy is

1173
01:12:26,440 --> 01:12:29,760
is is not, not what what what
gets progress.

1174
01:12:30,880 --> 01:12:33,640
I mean, we we see that the whole
progress in LLM due to open

1175
01:12:33,640 --> 01:12:35,440
source and then and so on and so
forth.

1176
01:12:35,440 --> 01:12:40,360
And I think this is also what
will, not what should, but what

1177
01:12:40,360 --> 01:12:45,280
will happen in in this space.
Yeah, I must say that I do

1178
01:12:45,280 --> 01:12:48,760
commend that you're releasing
the data and the models.

1179
01:12:48,760 --> 01:12:55,240
Open source is is a novelty.
And I think, yeah, it's good to

1180
01:12:55,240 --> 01:12:58,320
see because I like particularly
your paper now.

1181
01:12:58,480 --> 01:13:02,560
I mean, I think that has hurt a
little bit of trust in the CFD

1182
01:13:02,560 --> 01:13:06,200
world so far.
Instead, as you know, the

1183
01:13:06,200 --> 01:13:08,920
tradition and I'm don't sure
where the tradition came from.

1184
01:13:09,520 --> 01:13:13,440
If you develop a turbans model
or you develop a numerical

1185
01:13:13,440 --> 01:13:17,760
method, whatever it is, it's
almost always published.

1186
01:13:18,560 --> 01:13:23,040
It's very rare actually for a
model to make itself into a

1187
01:13:23,040 --> 01:13:26,920
large commercial, even, you
know, ISVCFD code and not to

1188
01:13:26,920 --> 01:13:28,280
have a corresponding
publication.

1189
01:13:28,800 --> 01:13:34,400
It's sort of seen as like a
thing that you don't make money

1190
01:13:34,440 --> 01:13:37,360
on the method, you make money on
the code.

1191
01:13:38,320 --> 01:13:43,840
That's sort of been the like
thing like open foam or star CCM

1192
01:13:43,840 --> 01:13:47,480
or anti fluid to others,
typically the models themselves.

1193
01:13:48,080 --> 01:13:51,840
You sort of know it's more the
the the little tricks that you

1194
01:13:51,840 --> 01:13:55,120
might do and but then more the
support and the code and all

1195
01:13:55,120 --> 01:13:58,640
that.
Whereas it felt like so far in

1196
01:13:58,640 --> 01:14:04,880
the ML4 CAE world, with the
exception of a a couple of

1197
01:14:04,880 --> 01:14:11,000
individuals or or companies,
most of the start-ups have not

1198
01:14:11,000 --> 01:14:14,760
published or made a thing of
saying, here's our exact method,

1199
01:14:15,600 --> 01:14:18,440
but we've implemented it in such
a way that, do you know what I

1200
01:14:18,440 --> 01:14:20,000
mean?
And I feel like that's hurt a

1201
01:14:20,000 --> 01:14:22,680
little bit of trust because it's
hard for someone to believe

1202
01:14:23,360 --> 01:14:27,040
results if they don't see a
publication about it.

1203
01:14:27,040 --> 01:14:30,680
So yeah, I don't know if I'm
alone on this, but I think your

1204
01:14:30,680 --> 01:14:36,360
strategy of publishing more, I
think actually helps at least

1205
01:14:36,360 --> 01:14:40,080
CFD people to feel a little bit
more like comfortable.

1206
01:14:40,600 --> 01:14:43,200
Which I, I, I think it's the,
it's the only way around.

1207
01:14:43,200 --> 01:14:46,440
And I if I make a bold
statement, if in the machine

1208
01:14:46,440 --> 01:14:49,320
learning world where things are
moving so fast you're afraid

1209
01:14:49,320 --> 01:14:52,360
that others are are copying your
stuff and overtaking you, you

1210
01:14:52,360 --> 01:14:55,840
should probably reconsider your
your capabilities as a company.

1211
01:14:57,640 --> 01:14:59,640
Yeah.
And the only thing I would say

1212
01:14:59,640 --> 01:15:02,240
though, which is why I've asked
you sometimes to up level the

1213
01:15:02,240 --> 01:15:06,360
explanations, is I feel maybe
even the biggest barrier is just

1214
01:15:06,360 --> 01:15:10,960
an understanding barrier that,
you know, if you're ACFD person,

1215
01:15:11,760 --> 01:15:14,280
you're probably going to read an
ML paper and be like, I don't

1216
01:15:14,280 --> 01:15:17,560
really understand this.
That's almost the challenge.

1217
01:15:17,560 --> 01:15:21,240
You know, it's like an education
thing is that the people making

1218
01:15:21,240 --> 01:15:23,680
decisions or higher up may not
even understand because it's

1219
01:15:23,680 --> 01:15:28,440
such a different field that that
is the challenge.

1220
01:15:28,440 --> 01:15:32,440
So anything that you can do to
make like, yeah, simpler

1221
01:15:32,440 --> 01:15:35,720
versions of papers or like
summaries of papers or, or

1222
01:15:35,720 --> 01:15:39,920
education, which I guess comes
back to why having a dual

1223
01:15:39,920 --> 01:15:43,440
affiliation between a university
and a start up is probably a

1224
01:15:43,440 --> 01:15:46,920
good thing.
Because I feel that you need to

1225
01:15:46,920 --> 01:15:51,320
be teaching and producing like
open source teaching material to

1226
01:15:51,320 --> 01:15:54,960
educate people as much as also
producing commercial solutions.

1227
01:15:56,400 --> 01:15:58,640
I don't know if you've noticed
that when you speak to customers

1228
01:15:58,640 --> 01:16:02,640
that sometimes they're just
knowledge of ML is part of the

1229
01:16:02,640 --> 01:16:05,400
blocker.
Yeah, but it's also the other

1230
01:16:05,400 --> 01:16:07,880
way around, right?
There's hardly any people who

1231
01:16:07,880 --> 01:16:10,720
have the knowledge of the real
domain expertise and knowledge

1232
01:16:10,720 --> 01:16:14,440
of ML because only if you have
the knowledge of domain

1233
01:16:14,440 --> 01:16:18,440
expertise and you know what ML
can do, you really find the

1234
01:16:18,440 --> 01:16:21,040
problems where ML can really
have an impact, right?

1235
01:16:21,520 --> 01:16:23,960
And, and and this goes in both
ways.

1236
01:16:23,960 --> 01:16:27,320
So education both ways is super
important and it's very, very

1237
01:16:27,320 --> 01:16:29,920
hard.
I mean, I, I see that myself all

1238
01:16:29,920 --> 01:16:32,840
the time that I can get lost in
details and, and, and think

1239
01:16:33,120 --> 01:16:34,880
things are obvious.
And on the other hand, when I

1240
01:16:34,880 --> 01:16:38,200
talk to real domain experts I
get get really lost and have to

1241
01:16:38,200 --> 01:16:41,720
ask questions and yeah.
Yeah, yeah.

1242
01:16:41,720 --> 01:16:45,920
Well, maybe This is why it still
will take a few years to reach a

1243
01:16:45,920 --> 01:16:49,760
greater maturity because both
sides need to upskill themselves

1244
01:16:50,120 --> 01:16:53,720
and, and, and sort of learn.
So great.

1245
01:16:53,720 --> 01:16:55,640
Well, really appreciate you
chatting.

1246
01:16:55,640 --> 01:16:59,880
I mean, what I'm going to do is
put some links in at least on

1247
01:16:59,880 --> 01:17:02,760
the YouTube to, to a couple of
the papers and the models that

1248
01:17:02,760 --> 01:17:05,280
you that we've referred to,
because I think people should

1249
01:17:05,280 --> 01:17:07,680
really spend a bit of time
looking at the papers.

1250
01:17:07,960 --> 01:17:10,400
I have a feeling that, you know,
we might need to chat again.

1251
01:17:10,640 --> 01:17:14,680
I think this field is moving so
quickly that everything that we,

1252
01:17:14,720 --> 01:17:16,880
it'll be interesting to see if
we listen to this conversation,

1253
01:17:16,880 --> 01:17:20,920
a year's time, are we going to
be completely wrong or in two

1254
01:17:20,920 --> 01:17:22,840
years time?
I very much hope so.

1255
01:17:22,840 --> 01:17:26,720
But what we won't be wrong is
that some things are moving fast

1256
01:17:26,840 --> 01:17:29,480
and that we will have fast
progress in two years.

1257
01:17:29,480 --> 01:17:32,800
So we will be, oh, we didn't
expect that to happen, so.

1258
01:17:32,800 --> 01:17:34,560
Yeah.
Yeah, that's going to happen.

1259
01:17:36,200 --> 01:17:38,400
That's a good thing though.
Maybe that maybe that's OK.

1260
01:17:38,400 --> 01:17:41,720
But yeah, for now, then, until
we speak again, thank you very

1261
01:17:41,720 --> 01:17:45,360
much and yeah, hope people that
enjoy looking at your paper and

1262
01:17:45,360 --> 01:17:47,120
your work.
Thank you very much, Neil for

1263
01:17:47,120 --> 01:17:47,480
having me.
