CDFAM Barcelona 2026 · Barcelona · 8 April 2026

HOOPS AI: Correlating CAD Geometry with Manufacturing and Business Process Information

Abstract

Advances in computational design and additive manufacturing have enabled increasingly complex geometry, but industrial AI adoption remains constrained by a fundamental challenge: relating CAD design data to manufacturing behavior and downstream business processes. Geometry, process data, and enterprise information are typically analyzed in isolation, making it difficult to understand how design decisions propagate across the product lifecycle.

HOOPS AI addresses this by transforming CAD and manufacturing data into unified, AI-ready representations built on the HOOPS platform. The approach focuses on extracting stable geometric and feature-level abstractions that serve as a common reference across engineering and non-engineering domains, without reliance on native CAD kernels or proprietary data models.

These representations enable systematic comparison, grouping, and retrieval of parts based on geometric and functional characteristics. The same representations can be associated with manufacturing signals, quality indicators, and production constraints, and extended to link with business process information in an IP-safe manner. This allows AI systems to reason about relationships between CAD geometry, manufacturing outcomes, and operational drivers.

Transcript

From YouTube’s automatic captions, lightly cleaned; expect some errors. Each timestamp opens the video at that moment.

Read the full transcript · 3,320 words

0:16 Hi everyone. Before starting I would like you to focus on this part and ask you the following question. What an expert would say about this part? A 20 years experience engineer within seconds could say whether or not this part is manufacturable. What will be the features driving the cost? What are the certain options and what tolerance has sense? And most of the time this kind of information is available in PLM system and other variants.

0:52 But sometimes the information is unstructured. And companies are very happy to have this experience engineer that can provide insight to the team within few minutes. So I have a couple important questions to ask. How can we connect geometry to that business knowledge? And more importantly, how can we spread the knowledge we have from one part to other similar? So that we can turn 3D shapes into actionable business knowledge.

1:24 I am Luis. I work for Tech Soft 3D. Where I hold the position of engineering manager for the machine learning team. I’ve been working in AI for the last couple of years and we have a mission on how can we use AI to add value in the additive industry. In today’s talk I will be showing you how geometry when connected to manufacturing business process becomes intelligence. I will be also showing how we develop Hoops AI that was the framework that we have been using for that.

1:57 We think that it so was in a good position to develop this framework because we are one of the companies and your preferred client resource for accessing data. We are capable of reading more than 30 file CAD formats and the same CAD and CA file formats. So, we have seen a lot of data over the over the time. And also we have technology for visualizing CAD data and also CA.

2:19 So, we moved the next step, how we can use AI for the CAD. But this doesn’t come with an easy answer. There are challenges when we want to address CAD data into AI. And the first thing I would like you to know is that CAD data is not a text. It’s not an image, it’s not a video. So, you LLM cannot write too many things with us.

2:46 So, we should index the CAD data to our brain. I agree, but in a suitable format. So, CAD data is multimodal. You have it the geometric information, the topology. You also have the PMI and other factors. Can be discrete and continuous. And even if you have solved that issue, how you can inject the data into your AI system, you need to solve the data set collection. You need a lot of data to make AI, specifically machine learning, to learn from.

3:17 So, maybe you will need to use a lot of file formats. How you handle large data sets that can also add value. How you explore that data. Now, finally, we know that we want to have a tech stack that is capable of handling all of that. But then we noticed that in the community there is ingredients. We have tools for opening and reading CAD files. And we have tools for using machine learning, like we have Python for instance.

3:44 So, the ingredients are there, the things is not really connected. So, we have created HOOPS AI on that mission. Our vision was to bridge CAD with machine learning so that we can unleash the next generation of applications. So, there are two main features in this framework. From one side, we have CAD data how we can access CAD data with transform this information into a machine learning on AI ready input.

4:11 And then at the left the right, sorry, how we can actually add some value. We need to solve some specific task. What kind of task do you want to solve? And since we are cooking provider and other things, we would like to try to solve the more general problem, not one simple one. There was something that we keep in mind when we were building this framework and that was very clear from day one.

4:32 We don’t want to focus on to building a model. We have to focus into building a factory because we believe like Open AI total doing, they are releasing different generations of model every 6 months and we believe the same with having this kind of trajectory. We will get into a model, it will get into a solution. We have a second version. So, the base is to have this factory that is capable of you to generating an advanced model.

4:57 And also important is that CAD data and your data is private. We cannot access. So, the tool needs to be ready for you to use on premise with your own data. And that was one of the key factors that we used to develop our Who Say I. And fortunately for solving that problem of how we can inject the CAD data, there is a scientific method actually. Hopefully, there is this geometric deep learning framework.

5:22 And this framework that is in the state of the art tells us that in order to inject the CAD data to machine learning input, we should transform the data into more representative vectors. We call that an encoding transformation. We can see here for instance in the middle what kind of data we extract as you can see graph for instance or some other more sophisticated metrics and and graphics so that we can get finalized with machine learning input that an AI can understand, a machine learning can understand.

5:55 And then when you have this, you want to focus now on when we implement this strategy in transformation and what we call so-so-called the ETL pipeline, we want to play the same for the CAD file. So, we transform the CAD into a machine learning input so that then you have a drive all those libraries that are very powerful to use and to actually learn something. So, the transformation in the middle, we just taking care of the validation.

6:19 We do it in an optimized way. We know we cannot allocate all that in memory. This all all is very clear and very transparent. We get our first data sets. In this video, I was showing the output of this first part of the presentation where is how we prepare this data, how you explore the the CAD file. Basically, you explore it by filtering. You will click on some kind of labels, on some kind of specificity, and you will try to visualize here.

6:49 By the way, I’m using the public data set halfway for this visualization. So, the user can easily query some part of the data so that we can visualize it. This are picture of but then I will select one of the files and then we will see the extraction that we mentioned before. Here, I am showing at the left what a mechanical engineer will see, a mechanical And in the middle, we have the face agency graph.

7:15 And at the right, we have also a point cloud visualization. This is a basic thing called a This is something known in literature. We filter the now, we focus on the real one, and we’re working on this one. And for each node, you can attach some attributes, areas, the local, and all of the extra information. You can see how by clicking on the node that is connected to the CAD.

7:36 So, I am basically translating that. The left is how a human would read the last two is how an AI will will do the race. And this is going just to do the first part. We have encoded our data and these are basic encoders. I’m going to move to a second example just to show you that these graphs become completely different because the graph grow with the with the with the face counts.

7:58 And also we have the the discretization. And even if these kind of encoders are very simple, we have also other type of encoders that goes to base on the a statistical geometric analysis. We have here for instance, I am we have need to solve this for a solving the task of machine feature recognition on that NIST CAD file. So we have to derive the face adjacency graph and you can see colors corresponding to the levels.

8:29 But also I want you to focus on the last two distributions, that is how you tell an AI how we can call neighborhoods. Basically, when you choose two random faces on the on the identity graph, you can have you walk between the political walk, sorry, between the two faces. You don’t have a notion of distance. And for motor mechanical engineer, you can see what is the a view factor maybe can resonates.

8:54 And basically this diagram tells an AI how far, how close, or how two faces to two pair of faces they look at each other. And the left down you can see that the two faces that I have chosen there, they are not touching each other because the distance doesn’t establish zero. But at the same time it’s telling me that the faces are not so far regarding to the full bonding box because they are mostly the left.

9:19 And at the same time the angle distribution is telling me these are not two coplanar faces because there is only one angle looking at them. So there is distributions angle, it means that the faces are regulars. So this is the kind of information that we encode to generate that in the the AI. So now that we have solved this first problem that how we know how we can get that out of the AI, we need to come back to our original statement.

9:40 What we can build shape intelligence based on that. So, the first thing we do to evaluate our framework is we use the previous architecture, something that is basically going to be implemented in the state of the art, and we solve these two problems that are are well known, a classification problem and a feature detection. Basically, this is supervised learning. You you label your data. You train your model.

10:08 And in the classification problem, for instance, the model tries to give you a prediction within the spectrum of the data he has seen. For instance, we train a data set with 45 by part labels. So, the AI will always give you a probability on that prediction. And at the right, we do the same thing with the faces, so we can predict manufacturing features. Whether these results were impressive and also building state of the art, there are a limitation.

10:36 This requires labeling data. And also, it requires solving one specific problem. If you need to change now the supervised task, you need to train you need to also solve the question the the problem again, and you also need to label our data. And labeling data is expensive. So, supervised learning requires labeled data. So, we have an idea I’ll say we need to solve the problem differently. What about we scale become that?

11:06 What about we teach the AI not to give a specific task, but to learn the language of shape? So, that we can escape without labeling. Now, the concept is very simple. The concept is something taken also from what is common. We take the data. We will build something in the middle that will generate the shape embedding with two premises. The shapes that are similar should be closer.

11:31 The shapes that are not similar should be followed. The intelligence that you put in the model in the middle is what basically will make the whole difference. And at the end the at the end the learned representation will help you to not use the level and then you will pre-pre-adapt your specific task. While supervised learning escapes with levels, unsupervised learning escapes with data and data that is available.

12:00 So, what kind of things they can unlock? They unlock a new challenge, a new kind of full value. Imagine that on the left you have one part and you manage to retrieve the five most similar in your database. The value of understanding shape now becomes finding similarity. But what happens when you have attached to this part all this information, manufacturing cost, the cost driving on the sourcing?

12:28 Then you move from understanding the shape to basically enable decisions. And you can measure be basically compare or spread information among the parts. The idea to based on that we have developed and it’s available in our tool kit something called oops embedding model. So, very simple model. We put something in the middle that we will pre-train with a lot of data so that these learned representations are are effectively learned, sorry.

12:57 And we these two primitives of finding similar shape that are close to different part are far are our strength. The The logic there is basically to learn from that. And the basic things we can do to prove that is that we took a training data set which is a probably one from Onshape. This is the largest CAD ABC data set. It is about 1 million of CAD files.

13:20 I don’t believe all of the part are labeled, but in In case we don’t need that labeling. So, we took all this data, as you can find this in our paper, and we pre-trained this hooks embedding model for the show. The training was done in a a few machine of 40 GB RAM, and it took more or less one week. But, it’s something that can be accelerated.

13:42 And once we have this pre-trained model on that public data set, we use a second data set, a second public data set, sorry, to test it. So, this data set is of 8,000 parts. It’s from our recent publication called the True Mechanical CAD Data Sets. This is the version two, by the way. And in this data set, they have some labels. In other words, we are going to remove that and find similarities.

14:04 And we’re going to test how the model perform on this new data set. So, but before doing that, I need to show the whole paragraph. How we find it the the similarity. So, we start from the data set. We use the pre-trained model, and we compute all the embeddings the vector. So, it’s one per part. So, in the end, we have 8,000 parts. We will index those vectors into a vector store that is optimized for retrieval.

14:34 In this example, we use Face, which is available in Python. This idea is coming from Rack. It’s just that instead of my embeddings being text embeddings, my embeddings are shapes embeddings. When we are doing that offline, then we will do the dynamic part, the online mode. We will take a query part part, we will do a query, so we will recompute the embedding models. It will be a vector.

14:59 And now, in that embeddings space, we retrieve the parts that are similar. So, the embedding vector of the query will go to that space, we’ll find closest neighbors, and those closest neighbors are going to to be given. Here, there’s an example. We have a query part. And I wrote I wrote a code that has an 8,000 parts. And these are the results that it finds most similar for the versions we have right now.

15:24 We can see some parts that are very similar or at least it’s but also there is something that we need to define here is that similarity is not objective proposition. So, we need to try to figure out what the model is really looking at. But here we can see also a similarity score and we see how the model has find almost the duplicate at the first square as it here is top and also part of the assembly.

15:53 I have a second example in here the data that has a lot of basically duplicated model, but that has basically slightly slightly changed a little bit. So, I can say that the first three part are mostly the same and we can see a difference in the last one of the industrial. So, the different comes that this feature that we can see anything orange it disappears in the first is different on the other two.

16:19 So, we can see that that’s what the model were looking when looking at the similarity. So, now what we can do with this? So, because this kind of similarity now is capable of working without supervision, you don’t need label data and you can also retrieve that. So, imagine that you I do the same query as before, but now this part has some information on the source. In some kind of an email, some kind of metadata, in some kind of features.

16:51 And there’s another part that has basically the information on the manufacturing process. And another might have code driver. The next step is to now build an engine on top of this that we can now insert information from the other similar parts based on what we know and what we have. What is important here is also that the shape similarity didn’t need all of the data to be trained.

17:16 And all the data that you will be appending day after day will still be available in this workflow because there is nothing being trained here after. And also what is important about the toolkit that the data can be trained on your data. Sorry, the model can be trained on your data. We don’t need to You need to reuse that pre-trained model so that you can have better results for your for your data.

17:37 Here also I’m going to be showing some of these future works. Actually we want to do the improve this model and have version two and a version three and duplicate what other companies are doing. And we want to add something what you call cosine definition similarity so that you can decide what is for you what are the two similar part are more likely for you. And we are also trying to see how we can improve this with a local rank.

18:05 For instance in this case I I might say that these three part that I have highlighted are more closer to the query than the second outfit you give me. And also in our journey of Victoria machine learning we would like to improve this kind of anomaly we see which is a very interesting anomaly. In our previous we mentioned that we wanted similar part to be close and not similar part to be pushed away.

18:29 What we notice in this anomaly is that we have two clusters because the three anomalies they look they’re similar to each other. So it means that there was something in training that make these two cluster part being close and when I’m doing the query then I got this kind of information. So we want to also add this on the planning of anomaly. Finally I would like to thank you all for your time and I’m showing here some other results.

18:56 You can find us in the coffee and please try the code so that you can have an overview of who we say I. And you can pre-train the model for yourself. Thank you so much. To learn more about the CDFAM Computational Design Symposium, access the archive of previous presentations, interviews with speakers, and information about future events around the world, visit CDFAM.com.

More from CDFAM Barcelona 2026

Digitizing Body-in-White Development with MAS Synera

Digitizing Body-in-White Development with MAS Synera

Juan de Dios Escribano Felguera; Tilman Steininger · SEAT; Synera

SubSimX: Interactive Subdivision-to-FEM for Computational Design

SubSimX: Interactive Subdivision-to-FEM for Computational Design

Johannes Müller-Römer · Fraunhofer IGD

Empowering Architects with Early-Stage Environmental Intelligence

Empowering Architects with Early-Stage Environmental Intelligence

Michele Pescatore; Carol Fanjul · AiA Life Designers

Computational Design of Personalized CPAP Masks

Computational Design of Personalized CPAP Masks

Anne Pasman; Emmy Kerssen · Saxion Hogeschool

Raven AI: The Future Is In The Spaghetti

Raven AI: The Future Is In The Spaghetti

Moritz Rietschel · Raven

Redefining mechanical engineering in the age of AI

Redefining mechanical engineering in the age of AI

Rhushik Matroja · Cognitive Design Systems

Bridging Data to Geometry with Implicit Modeling

Bridging Data to Geometry with Implicit Modeling

Wesley Essink · Siemens Digital Industries Software

Architected Porosity Informed by Real-World Data for More-Than-Human Thermal Comfort

Architected Porosity Informed by Real-World Data for More-Than-Human Thermal Comfort

Maria Claudia Valverde Rojas · University of Stuttgart, IntCDC

Real-time Multi-Physics Collaboration for Real-world Engineering

Real-time Multi-Physics Collaboration for Real-world Engineering

Nikolas Borrel Jensen; Oliver Littlewood · Pasteur Labs

Register for Updates and Discounts on CDFAM events.