CDFAM CD/DC 26 · Washington DC · 15 July 2026
Measuring Shape Fidelity in Generative CAD Models
Abstract
Generative AI is rapidly expanding what designers and engineers can create in 3D, but visual plausibility is not the same as geometric fidelity. This presentation asks a practical validation question: how close are AI-generated CAD models really to a target desired shape?
We introduce a feature-vector workflow using the Metafold Shape Similarity technology to compare generated models against reference targets. Each model is encoded into a geometric feature vector, enabling direct comparison through aggregate similarity scores, coordinate-level distance ribbons, scale-normalized metrics, and side-by-side 3D previews. The result is a repeatable method for moving beyond “looks right” evaluation toward measurable shape correspondence.
Using examples from current 3D generative design workflows, the talk demonstrates how feature vectors can expose where a generated model preserves intent, where it drifts, and which geometric features contribute most to the gap.
This approach offers a lightweight validation layer for AI-assisted CAD: fast enough for iteration, interpretable enough for engineering review, and concrete enough to support model benchmarking.
Transcript
From the speaker’s corrected captions. Each timestamp opens the video at that moment.
Read the full transcript · 3,331 words
All right. It’s great to great to be back at CDFAM. Always a really special conference. I always get to chat with really really awesome people. So, today I’m talking about shape fidelity and I think the the kind of tenor of this talk is a little bit different. It’s a little technical. I’m a mathematician by background. I did learn pretty early on that I was not a great research mathematician.
0:29 Spoiler, it’s was very hard. But but I’m, you know, not not a terrible applied mathematician. And so, I got really interested in in all things to do with geometry processing and and manufacturing. So today we’re going to talk about measuring similar similarity between 3D shapes in the context of 3D gen AI. And before I kind of do this, how many people are kind of geometry processing, computational design folks in this room?
1:25 Okay, maybe 35%. Any mechanical engineers? Okay, this will be very interesting. Okay. So let’s get into it. So we are talking about this in the context of AI generation for 3D models and why do this at all? Well a couple a couple reasons. There is a pretty huge market for this. Even long before AI 3D interfaces are known to be difficult to learn, hard to master just they’re just it’s it’s very difficult still to you know create good 3D models.
2:05 And if we could really you know democratize how people create 3D data, good 3D data, then I think we’d be able to you know build the things that we need to build. And we need to build we do as a as a civilization we do need to build things. We need to build them quickly. So specifically there are some really high value use cases for 3D AI in similarity and geometrical search.
2:31 Feature recognition costing and quotation is a really big one. And how much does something cost? If you give me a part, I need to give you a price. What should that price be? And and so the the other part of this is that there’s been a kind of explosion of interest in 3D AI. If your LinkedIn feeds are anything like mine, they might be completely inundated with the latest like 3D gen AI.
2:56 And weirdly, not so much the latest 3D models, but but the benchmarks. Assoc associated with them. So that’s what this talk is. It’s it’s about looking at the not necessarily the models that create these for for 3D genai, but the benchmarks that measure how good they are. And so you know my my my team’s contribution to this is to improve the quality of results that you can get out of out of a 3D AI /ML pipeline.
3:33 And I always like to put the slash ML cuz I’m still like to do that. Okay. And before we get into the the to the next bit too, the question is, well, if we can create 3D models using AI, are they useful? And I I think it’s really important to kind of just cover these quickly. What does useful mean in this context? This is not an exhaustive list.
4:01 This is the list that I typically work off of, but it should take less time to create that model than with other tools. The generated part that you get should be within some specification and a small amount of user input or information is required to generate the part. So these are kind of like the things I look for in a 3D genai pipeline. In contrast, things that are not useful.
4:27 It’s not useful if the generation process takes hours of iteration because I could, you know, go do this in some other tool and and do it probably better. It’s not useful if the the generated part fails to meet a target specification. This, you know, this assumes kind of parts in the engineering and manufacturing world. So, it’s, you know, game assets are maybe a different a different beast.
4:49 And this last one I think is Yeah. So users shouldn’t need to provide comprehensive images, tests, drawings, and other spec other context to generate this thing. If you’re giving the model like an enormous amount of information to generate it, you’ve not really gained anything. It’s kind of transforming that knowledge and context into a 3D form. That’s cool and could be interesting, but you’re you can’t I think in that sense you you can’t say that you’re really generating three a new 3D model.
So yeah, these are these are the the things that that I look for in 3D pipelines. So maybe to point out the real issue at the heart of all this is that AI has a 3D blind spot. This is due to a number of factors. 3D data sets are limited and proprietary. Metadata, CAD history, named features can be incomplete or wrong. Mechanical engineers, would you agree that this can be true or are all your metadata fields 100% present and correct all the time?
5:57 Probably not. I’m going to assume not. And and so AI models for text and images, which work astoundingly well, they just don’t extend to 3D. And unlike you know text and images there’s no canonical representation of a 3D shape. This sort of the way you represent a shape in 3D is this additional layer of complexity between what you might want to learn what what you want these AI systems to learn.
6:36 So, a little bit of historical context. At Metafold, we did a bunch of work for a very large ondemand manufacturer. This happened over the course of 24 and 25. And so we had been really thinking about this problem very differently. This this this question of of pricing, predictive quotation, stuff like this. And over the course of that project we deployed what we call the Metafold embeddings. This this approach to to encoding shapes.
7:13 We we deployed the embeddings into this on demand manufacturer and found that they outperformed so in 6 months they outperformed the the the feature vector that they had developed over the course of 10 years to support their their their their cost quotation engine and you know so that’s nice we can give ourselves a pat on the back but but I think it’s very interesting to ask why is why is that true?
7:38 How did how did that happen? And so that project the kind of deep amount of work that we did for that in conjunction with this like explosion of interest in in benchmarking the outcomes of of Gen AI 3D geni is is I think kind of why I’m here today. We felt like we needed to say something about this. So yeah, there there was kind of this this crazy thing where in May six benchmarks were produced to measure how effectively AI can produce 3D shapes, not just models, but like different literal yard sticks to to measure the success of these things.
8:24 And then there’s more there was more last month. I saw one like yesterday. So, for a field that was kind of not in the in the spotlight, it has really risen to to kind of, real prominence. And and these are great benchmarks. And so we spent a lot of time you know reviewing what went on in the methodology here and trying to figure out how they compare with the work that we have done.
9:01 And the point is that a kind of concerning thing happened with these benchmarks, which is they all use the same ways to measure similarity in 3D shapes. And these these methods to measure shapes miss a ton. And so the I think the effort to benchmark things is really valuable. But I you know here today to try and like offer a course correction to to how we measure the differences between shapes.
9:31 So why is it hard to to measure I I keep using words like similarity or distance? First of all what do I mean by this and why is it hard? So what I mean by this is if you give me two shapes like two 3D models I need to give you a number back and that number is going to be the distance between them. And so this is easy to do if you give me two points, right?
9:53 I just measure a line, I give you back the distance. And but but it’s it’s quite difficult to do this properly using 3D shapes. If you’re in computational design, you’ll probably reach for a handful of common ways to do this. You the first thing you might think of is like a bounding box. I’ll take two shapes and I’ll put their bounding boxes around them. That’s, you know, six coordinates or eight coordinates and I can measure the distance between those.
10:23 Or I could do some kind of point cloud comparison. The problem is that when you do this, you you kind of if you if you use these standard methods, you’re going to get distances that are either strongly biased towards zero, as in not at all similar, or strongly biased towards one, as in it’s the same shape. And so if I actually overlay this histogram of the outputs of similarities from the benchmarks.
10:53 So this is just numbers from the published benchmarks. You’ll see that exactly this happened. Those numbers are all piled up around zero and they’re and a few of them are p piled up around one. There’s many more zeros than ones because they’re not a lot of identical shapes. And so the green region here is what is happening in the middle. And that’s that’s where like the good stuff is.
11:15 Yes, we want to be able to tell if something’s identical, but we really want to know, especially for the purposes of AI, what’s happening in that that nice middle ground. So, in comparison, we ran and this is sort of the theme of this study. We sort of ran the benchmark methodology and then we ran the Metafold embedding methodology. And in contrast with the Metafold embeddings, you see we get a really nice distribution of distances measuring interesting things we hope in in that middle ground.
11:52 So armed with this ability to ability to recreate the benchmarks and the Metafold embedding methodology, we did a couple studies on the data sets from the benchmarks. And the first question we were trying to answer is do the benchmarks the way that they measure similarity does that capture engineering intent. So we looked at the Autodesk Fusion data set. We also handcrafted some defeatured models and ran both checks in parallel.
12:17 I should also just introduce the Metafold embeddings very briefly here. This is a little web app that you can play around with where you upload a bunch of shapes and it computes the embedding. That’s one of the columns in there. And you can just measure the similarity and it gives you this nice barcode, this heat map of how these shapes are different. What exactly that’s measuring is not particularly interpretable and it’s that’s by design for for math reasons, not like IP reasons.
12:58 Maybe also IP reasons, but but here’s a consequence of of these embeddings. So, here are two shapes in a sequence of design moves that an engineer might do. You have this sort of tube thing and you’re trying to make I think you’re trying to cut a slot in here. So, you do this, you you apply like a profile and you cut a slot in here. The this represents an a big change in design intent right on before the slot you have something you can mill and like you could turn it you could probably buy stock in that shape.
13:34 But on the right you have like a very interesting profile. Like someone has to go figure out how to how to manufacture that. And as a consequence of how those benchmarks measure distance, it doesn’t it doesn’t capture this at all. It just thinks they’re basically the same. There’s a little variance. The Metafold embedding drops considerably. And same for this. You have a a plate with on the left it’s you know a standard sheet of some thickness with some holes drilled out.
14:05 This is very simple to do. And on the right you have these tabs and as soon as you add these tabs you you’re it’s not a it’s not a sheet metal part anymore. You have to figure out how those tabs go on. So again there’s a dramatic change in the intent between these two shapes. And again, purely because of the choice of metrics from these benchmarks, it doesn’t pick up on this at all.
14:34 It thinks they they are basically the same when they’re they’re very much not. And so we can we can go further in our analysis here and look at this big matrix of comparison between a handful of shapes from the benchmarks and the similarities similarity scores from Metafold. And the way to read this is on the diagonal you’re comparing the same shapes. So that should all be one.
15:00 They’re identical. Distance is well in this case one similarity is one. And then it’s blue when similarity is zero. And on the right you can see the benchmark scores just drop off. They measure identicalness but nothing else. And then on the left hand side you can see that the the the metaphor distribution has a much richer variance. This is the same calculation but applied to the design intense study.
15:26 Same problem clustering around highly identical parts. And then here you can see that it’s the Metafold distribution has a has a much richer distribution here. So frankly I think do the do the benchmarks capture design intent? No, they don’t. Cuz they they they can tell you if a part’s the same or they can tell you if it’s not at all the same. But they miss very critical engineering features.
I think it was maybe about 10 years ago that a lot of geometry processing folks decided that it was not enough to study these algorithms kind of in the lab where you can control the the health of these models. It’s important to run studies like in the wild. This term kind of in the wild is usually associated with a lot of geometry processing moves. So we wanted to do that as well.
16:22 So what happens can you run these benchmarks? Can you robustly run these benchmarks in the wild? If I if I have a bunch of parts lying around, can I run these benchmarks? How how long do they take to run? Do they take to run? What happens when they disagree with the Metafold embeddings? And so we want to look for, you know, completion timings, things like that. We have a a data set we call the Metafold ABC 1000.
16:44 If you do any data science in in 3D, you’ll have come across the ABC data set. It is frankly kind of a terrible data set because it has a huge amount of bias in parts. And so the Metafold ABC 1000 is is a nicely distributed set of 10,00 parts so that you can, you know, test your your latest and greatest geometry processing algorithms. So we wanted to run this.
17:11 We we use it all the time internally for benchmarking and we wanted to run the benchmarks and compare with with Metafold. And this failed the the the benchmarks did not complete they did not complete because of meshing errors. They in the documentation also it’s it usually says that they require all these shapes to be completely watertight and so on. And so in the wild you will not get this condition and so they failed to fail to compute.
17:46 Also expected compute time was 150 days. This is not acceptable. And these registration based methods this is kind of in the weeds but this is one of the fundamental ways they do this. They just don’t work work well in these nonidentical situations. The medical embeddings were happy to do this. We you know completed completed the embedding in about 20 minutes got a nice distribution. We did want to kind of do a bit of a disagreement study.
18:22 So we looked at where the scores assigned by Metafold and the benchmarks differ a lot or where they align and where they differ. And what we found probably no surprise is along the same theme where there is a clear match for finding a similar part in the data set both both find it if if they can compute it. They both find it but in cases where there’s not a good match the embeddings give a much more defensible answer.
18:52 And we also do this at a tiny fraction of the computational cost. And to kind of dig a bit deeper into the performance comparisons, we we looked at the the benchmark performance and over the over a handful of of shapes that it could compute. And on average, it takes about 30 to 45 seconds to index to to compute the similarity between two shapes. Whereas Metafold takes about you know 2 to 3 seconds per shape and then in terms of querying similarity it’s down at the like you know 40 to 50 millisecond range.
19:43 The other point too on on performance is on the topic of compression. So the Metafold embedding methodology takes a CAD file of arbitrary complexity. So it could be 20 megabytes, it could be a gigabyte, 2 GB and reduces it down to something like 2 to 300 numbers. So that’s an enormous compression ratio and you can’t recreate that file. It’s gone. But you can do things like data science, look for trends, clustering, you can train AI models using it.
20:19 So there’s an absolutely enormous compression ratio. And in contrast, the benchmarks require you to keep all your CAD files like around. They require if you want to check the distance between two shapes, you got to go look up those CAD files, do something with them, and then compute the shape like through registration or something. And so this basically provides no compression. And comp compression is crucial for providing fast inference and just you know scaling this up to the data sets that that it needs to be scaled up to.
20:59 Okay. So a few conclusions here to zoom out a little bit. There’s a couple things you know if you can walk away with a a few conclusions here. This is one of them is that simply how you measure similarity and remember similarity is sort of synonymous with the distance between things and if you’re all mathematically inclined when you hear distance that’s a key part of reinforcement learning and derivatives and stuff like that.
21:23 So there’s kind of a fundamental math thing going on here but how you measure this has a direct impact on the quality of u both benchmarking generative 3D AI but also on on the on the models themselves. A little bit of a sharper conclusion I guess which should be no surprise at this point. A robust scalable embedding methodology is critical to to scaling this up to to usefulness I would say in terms of training the types of models on the on the sizes of data set data sets that we need to train them.
22:05 You just need this compression amount and you need the fast querying otherwise we won’t get there before you know the heat death of of the universe. That’s it. I’m happy to chat with all of you about this in more detail. But thanks for your time.
More from Daniel Hambleton
More from CDFAM CD/DC 26

Agentic Engineering: Generative AI in structural applications
Sergey Pigach · CORE studio | Thornton Tomasetti

When Failure Is Not an Option: Bringing Certifiable AI to Engineering Design
Rhushik Matroja · Cognitive Design Systems

From Requirements to Manufacturable Systems: Agentic AI on a Live Engineering Knowledge Graph
Chris Helmerich · Celedon Solutions

The Digital Thread In The Real World: Multiple Partners, Multiple Tools, One Truth
Austin Herrema · Istari Digital








