CDFAM Barcelona 2026 · Barcelona · 9 April 2026

An Engineer’s Approach to Integrating Machine Learning in Generative Design Tools

Abstract

Machine learning and Artificial Intelligence have enormous potential as tools for generative design in engineering. However, most industry efforts remain stuck in research prototypes, brittle bespoke models, or disconnected add-ons that rarely survive real engineering workflows. In this talk, we will present an engineer’s approach to integrating machine learning directly into production-ready generative design tools, drawing on our experience building the fastest physics-driven thermo-fluid optimization platform on the market. Rather than replacing physics with opaque black boxes, our methodology uses ML only where it strengthens engineering outcomes.

I will show how ToffeeX’s ML developments accelerate design exploration and automation while preserving full control of engineering intent, seamlessly extending our existing topology optimization engine which is already used daily in real production environments. This talk highlights why our approach, built on smart algorithmic design rather than brute-force model training, achieves the reliability, manufacturability, and speed required for real-world engineering.

Transcript

From YouTube’s automatic captions, lightly cleaned; expect some errors. Each timestamp opens the video at that moment.

Read the full transcript · 3,502 words

0:17 Hi everyone. My name’s Thomas obviously I’m from ToffeeX. That’s the title of my talk. This is kind of a bit of a hodgepodge of like my thought about applying machine learning in generative design after working in the development of engineering software for about 10 years now. Last no cool first I’ll talk a little bit about what Coffee X is, who we are for those of you who don’t already know.

0:45 Although we’ve been around for almost 6 years now. So we’re a generative design software tool based on topology optimization primarily. But we specialize in thermofluids so things like heat exchangers, cold plates, manifolds, valves, all this kind of good stuff. And we’re cloud based and we’re really fast and yeah it’s it’s it’s physics driven. So underlying it all there’s a CFD solver. So, if you go buy toffee at the shop, right now, there’s no like, surrogate modeling or anything like that.

1:20 It’s all based on good old-fashioned from the ‘7s finite volume CFD. And thus, that’s how big we are. We’re based in central London. We’re spin out from Imperial College London. We work with all these people. So, primarily we work in aerospace, energy, automotive, and, nuclear, all sorts. We work with a lot of companies. There we go. All right so what is topology optimization? What is multifysics topology optimization?

1:49 So the principle is you start with a design domain. So basically where can I build stuff? Then you kind of set up your physics problem. So you set up your CFD problem. So what is my flow rate? What are my heat fluxes? And so on. That’s what it looks like. You describe you make a mesh. Then you solve your CFB. You saw your topology optimization algorithm. You update the design on and on and on and on.

2:15 Yes, it has converged. This is what it looks like sped up like 5,000 times. So what you can imagine here is you have a coolant coming through. You’ve got hot walls. You’re trying to cool the hot walls as much as possible. And the optimal design is a series of pin of fins that cross recirculation without separation. So you get high convective heat transfer and there’s no turbulence to cause losses.

2:44 Although obviously in some cases it might generate turbulent generating structures because turbulence is also really good for heat transfer. That’s what it looks like. So that’s kind of the underlying technology. But what we quickly discovered was that you know most engineers don’t like black boxes where you kind of like give a bunch of inputs and you press go and then you go away and you come back 2 hours later and you got a design with no explanation how you got there.

3:08 So primarily what we’ve been developing over the last several years is controllability over the topology optimization. And this doesn’t just mean manufacturing constraints doesn’t just mean we we satisfy some overhang angle for additive or or or or some some some radius for for for a stamping process. It means capturing engineering design intent saying so engineers can say I need to achieve this and this is how I want to achieve that.

3:35 I want to achieve it with a lattice. I want to achieve it with some sort of existing structure that I can only change by some percentage. I want to achieve it by maintaining these regions to be static to be unchangeable. And so it’s kind of like bringing into this kind of like CAD based design intent while also in the regions where we want to allow design freedom to allow basically the topology optimization algorithm to work.

4:03 And like we’re used like I said by all those companies. So this this if people were here last year, you know all these. But this is a this is a liquid natural gas vaporizer installed in several natural gas terminals in Northern Europe. This is a piece of high pressure die casting tooling, one of several installed on Toyota production lines. And this is a coal plate for an AC/DC inverter which may go on well hopefully, fingers crossed, is going to go on a German car in 2028.

4:38 And that’s why it’s called Coffee X. Okay. So, machine learning, okay, I did not look at the time when we started, so let go. So, we we’re pretty quick as a topology optimization tool. We’re not as quick as like some sort of native machine learning or AI based tool. And by the way, everything I say now, I’m going to be talking about physics surrogates rather than agents.

5:01 Agents is a whole separate thing. And I’m speaking in rapid in Boston next week about agents. So come to that if you want to know what I think about agents. But so people in the room working with surrogate models probably recognize these. It’s certainly very true at Toffee X, right? So there’s three problems, right? So one, you try to develop something, it works really well and it’s a research prototype and you got really good benchmarks and it gets really good results.

5:26 But that’s the only thing it’s good at and then it breaks or or or whatever. So pins like pins pins seem very promising about 5 years ago. Has anyone tried to make a good pin recently? No. Cool. So or you make a bespoke model which is really good at making pumps. So like the talk from Sim Scale that’s probably really good but it’s only good at making pumps.

5:51 And with that’s not a problem, but how do you then scale that to different size pumps, for example? And also uncertainty bounds. I’ll talk a little bit about this a little bit later. And then another add-on, right? So, how do you how do you get people to use it rather than develop it, put it in your product, and then cross your fingers? Yeah, agents. Okay. And many kids, you get three for the price of one.

6:16 So, I don’t know why all my formatting’s gone wrong now, but there you go. So when we’re working on machine learning at TFS, these are kind of the questions I’ve sort of narrowed down to. If we can answer yes to all these four all four of these questions and it probably will work. So yeah, I mean you can read it, right? So is there there’s is there a a physical approximation a well characterized physical approximation that we can address?

6:45 Can we generate training data? Okay, we talk a lot about data, but for fluids, which is what we work in, you need a lot of data, right? So, you saw it seems there was like thousands of simulations to get that that that pump working, that pump model working. Are the failure mills physically interpretable? Again, I’ll show you some pictures and I think that that’s that would be the most convincing thing for you guys.

And can your engineers work without without having a PhD in machine learning? So, here are some projects we’ve worked on at Toffee X. So this one I should put I should have put it on here. This was funded by the UK Innovate UK about several years ago and and our idea was like we do a lot of internal flows in heat exchanges and so on and conventional RANs modeling for fluid flows is pretty garbage at modeling turbulence in internal flows.

So you you get these huge kind of errors associated with doing CFD through a heat exchanger. Unless you’re doing large eddy simulation or DNS or I guess ladder boltsman in some cases as well. But most people in industry don’t do that. So they want to use RANs. But we thought well rand models basically re they’re physically sound. There’s a couple of questionable approximations but you know generally they’re physically sound and they rely on these experiment empirically generated coefficients.

8:07 But these empirical coefficients are generated on a subset of data and you need to try and basically tailor these to to to your problem and they’re really bad at sort of predicting separated fields. So this is this is kind of like a periodic thing here and you can oh okay now it works again. You see these things here these are like big separations. And so if you do a rand model it’s pretty bad.

8:33 And what we wanted to do was basically say, okay, can we find better coefficients for the trans model based on LE and DNS simulations? So, we don’t want to replace the RAM. We don’t want to wholesale replace CFD. We just want to be able to make it better so we can get lees or DNS quality results at the cost of a RAM simulation. And so here you go.

9:02 Oh, this is all messed up again as well. Interesting. H. Okay. So don’t know what’s going on here, but okay. So, this this was the le this was the ales here. This was no, this was the this was the RANs. Sorry. This was the machine learning and this is the alse flipped, but you can see essentially you get much closer compared to the RANs here as well.

9:28 And so what you’re doing is for these internal flow geometries, you can run a simulation which takes as long as a rans and you get the same fidelity as an ales which is a speed up of like 50% 50 times not 50% 50 times. So another project we we we’re working on we worked on this is actually done by an intern is completely insane that she did this all in 3 months.

9:50 But basically to build a machine learning surrogate of the flow through a heat exchanger which is okay people have done this fine whatever but the idea is then if we can use this machine learning surrogate of of the flow through a heat exchanger can we then use that as a boundary condition for a topology optimization. So the manifolds delivering fluid to a heat exchanger obviously have a big effect on the performance of the heat exchanger cuz you want to sort of get a uniform flow across the heat exchanger and maybe get a higher flow in regions where there’s larger delta t and so on and so but if you’re designing you know a manifold using topology optimization you cannot kind of like predict if I change this in my manifold how will the flow change in my heat exchanger without some sort of model for that heat exchanger.

So basically that’s what we did. And we trained again on well we trained on three different types of data rand ales and DNS. And the idea was we do loads of rands because we can do loads of rands and then we do lower numbers of also expensive. And we can basically assign importance to these results. So basically these weights basically say to the to the to the regression we’re building this is most accurate this is middle accurate and this is least accurate.

11:00 And what we found essentially so this is basically if if all these dots were on this all these dots were on this dash line you get a perfect prediction and what happens around here is basically the wheels fall off and that’s basically where you go from from a laminer flow through through the transition to turbulence and the nonline nonlinearity of the nav sto really takes over. And what that means is you need heaps more data in between here to get these all to land on the on the line.

11:33 Interesting. It’s also where our BNS data ran out, but I think that’s just a coincidence. Heat transfer predictions. They’re all there. They’re equivalent essentially to the pressure loss predictions. Last thing, so I talked a little bit about engineering designs and I talked a little bit about, manufacturing constraints on the topology optimization. So one manufacturing constraint is the overhang angle for for additive manufacturing. And this is quite an expensive computation.

12:01 You need to first calculate basically where you’re where where you’re violating it and then how to repair it in a way that is least disruptive to the fluid flow. And then you want to try and incorporate that into your topology optimization loop. Throughout every iteration of the topology you can basically find the correction to satisfy your self-supporting overhangle overhang angle overhangle that’s good so there’s we I’m not going to tell you our our approach but it basically it’s it’s it’s a deterministic algorithmic approach and it’s it’s a little expensive but it it it works really well this and and it gets called thousands times now but It depends on three three parameters and it depends nonlinearly on these three parameters.

12:51 And so when we basically were trying to figure out what are the optimal parameters to use in our our in our topology optimization algorithm, we got kind of lost. And kind of what happens is you start running it and it kind of gets stuck here. And so you can see these steps these steps are these are just iterations. But yeah, it gets stuck and it stops moving.

13:09 And if you choose an optimal set of parameters, basically it finishes it real quick like that bang. And so basically you go from like thousands or tens of thousands of steps from the top one to to be able to complete this to basically 400. And all of a sudden it’s tractable to put this into the topology optimization algorithm. And basically all our regression is doing is calculating what are the optimal parameters for my overhang correction for a given mesh.

13:39 And so we’re accelerating heaps with machine learning but probably not in the way you’re thinking and suddenly this is now very tractable in terology optimization. So what do they have in common? I have no idea where I am on time but yeah we had a known target that we wanted to do that was physically that we could physically represent whether it was like surrogate modeling or some sort of representation of the overhang.

14:07 We had high fidelity data we could generate at a reasonable cost and we’ll talk about cost in a little bit. Failure modes were kind of clear. So in the turbulence we could bound the coefficients that the model generates. Clearly it’s we can check the overhang angle. We know like what is the correct pressure loss and not and also the engineer can look at the output. So again the engineer can look at the coefficients predicted by the turbulence and say this is a weird coefficient that may or may that might not be correct.

14:42 And so on. Okay. So now zoom out kind of like what’s going on. In in in in in the industry right now. So what why why cuz this is maybe I should also say like when I speak to our customers and we or to our prospects even and we go, “Hey, we got this really great tool. It’s going to save you so much money.” no one says, “hm, but can it do AI?” Most of the time most of the time they they just care if it works.

And so the the it’s it’s a question kind of like if there’s so much like interest and like to be clear like we’re super interested in building these surrogates as well because obviously if we can take our topology optimization time from like 5 hours to 5 minutes that’d be incredible. And what we found first of all there is a data problem like especially for simulation data simulation data is like hard one IP for these large engineering companies.

15:40 They’re not they’re not going to share it with anyone else. You have to work with the simulation data within a single organization. Or you can generate your own. And if you want to generate your own this is from a paper everyone probably already knows about this paper, but yeah, if you want to generate your own, you’re going to be spending hundreds of millions just to generate data before you even trained your model.

15:59 At top, we don’t have $100 million. So there’s this one as well, which I think this is the low fidelity ceiling, which is kind of like what I was talking about earlier about like RANs versus ales in the context of fluids. So it’s not just the difference between low fidelity and high fidelity simulation tools. It’s within a single RANS tool. There’s a huge amount of variation. This is a graph from a NASA drag Oh, sorry, AAA drive prediction workshop.

16:27 It’s on the NASA website, which is why I said that. But they did this in like 2000 or something like that basically and used like a CFD tools to basically characterize the drag around an air foil and they evaluated like 270 CFD tools. So this is just 30 tools. And this is the variation in the drag coefficient around this air foil predicted by each tool. And by the way, the average here, this black line is actually above the experimental measurement.

16:53 So when you’re training on CFD tools, you’re training already on questionable data that is often not backed up by experiment. So if you have a 7% error here, which is what this is, and then your machine learning model has a 7% error on top of that 7% error, at what point does it become unacceptable for your surrogate? And then there’s also like the the the interpretability problem.

So here’s two results. One of them is from a simulation tool. The other one is from a machine learning sort of guess trained on trained on that. Which is the correct one? This is same Reynold number, same everything. It should be interpreted by an engineer. Which is the correct one? Right. This is this is turbulence. This is this has a huge effect on the pressure loss across your your your system.

17:44 And I’ve blacked out here what’s going on here. But you see the scale the same. I wouldn’t be able to tell. I doubt any other engineer would be able to tell. So if your machine learning model is giving you this difference, how do you know it’s correct? You got to run the CFD. So just run the CFD in the first place. Here, same sort of question. This is the same thing.

18:03 I’ve inserted another one there as well. But yeah and there’s this other like interesting paper I found published at ML for PS last year which kind of like tried to build a a a formal proof of like why neural networks are pretty poor at extrapolation be beyond their training data. Not sure I buy it completely but it’s it’s an interesting paper. I encourage you to check it out.

18:29 But yeah again which one’s correct. Finally regulatory void. So we work a lot with aerospace nuclear energy and so on. So we have this our customers have this to to to problem to work with as well. There’s no regulatory framework for any of these engineering fields about the use of machine learning tools. They have their surrogate model I guess like frameworks and I’ve used them in the past in in pre in a previous life.

18:58 But if you’re trying to build I guess what we’re calling foundation models now where there is no very clearly defined sort of like area where you work there’s no regulatory framework for this and like yeah this is this is kind of like think okay it’s really cool if you use it in Formula 1 but but I watched that Chernobyl miniseries last week again and yeah okay so finally to wrap up this is kind of like where we sit at this with X which is kind of like what are actually trying to do here basically?

19:29 Are are we are we are we still trying to do physical modeling based on like the fundamental laws of physics just conservation of mass, momentum and energy? In which case let’s just solve them equations cuz they capture it all perfectly. And then when we do like we we need to know when not to trust it. So like I mean the these pictures like I yeah okay and then agent AI.

19:53 Yeah, I I got lots of thoughts about agents. I’m much more positive about agents. But come to my talk in Boston in Boston next week. All right. Thank you. To learn more about the CDFM computational design symposium, access the archive of previous presentations, interviews with speakers, and information about future events around the world, visit CDFAM.com. Come.

More from CDFAM Barcelona 2026

Digitizing Body-in-White Development with MAS Synera

Digitizing Body-in-White Development with MAS Synera

Juan de Dios Escribano Felguera; Tilman Steininger · SEAT; Synera

SubSimX: Interactive Subdivision-to-FEM for Computational Design

SubSimX: Interactive Subdivision-to-FEM for Computational Design

Johannes Müller-Römer · Fraunhofer IGD

Empowering Architects with Early-Stage Environmental Intelligence

Empowering Architects with Early-Stage Environmental Intelligence

Michele Pescatore; Carol Fanjul · AiA Life Designers

Computational Design of Personalized CPAP Masks

Computational Design of Personalized CPAP Masks

Anne Pasman; Emmy Kerssen · Saxion Hogeschool

Raven AI: The Future Is In The Spaghetti

Raven AI: The Future Is In The Spaghetti

Moritz Rietschel · Raven

Redefining mechanical engineering in the age of AI

Redefining mechanical engineering in the age of AI

Rhushik Matroja · Cognitive Design Systems

Bridging Data to Geometry with Implicit Modeling

Bridging Data to Geometry with Implicit Modeling

Wesley Essink · Siemens Digital Industries Software

Architected Porosity Informed by Real-World Data for More-Than-Human Thermal Comfort

Architected Porosity Informed by Real-World Data for More-Than-Human Thermal Comfort

Maria Claudia Valverde Rojas · University of Stuttgart, IntCDC

Real-time Multi-Physics Collaboration for Real-world Engineering

Real-time Multi-Physics Collaboration for Real-world Engineering

Nikolas Borrel Jensen; Oliver Littlewood · Pasteur Labs

Register for Updates and Discounts on CDFAM events.