CDFAM NYC 2024 · New York · 2–3 October 2024

State of the Art B-Rep Generation

Abstract

Boundary representation (B-rep) 3D models are the standard 3D representation used in the manufacturing industry. However, only recently has machine learning research begun to make progress on generative models capable of producing B-rep models. This talk will give a summary of the current state of the art for generating B-rep models. In particular it will cover, BrepGen, our recent work using diffusion models, that have proved extremely successful in the image domain, to the problem of B-rep generation.

Transcript

From YouTube’s automatic captions, lightly cleaned; expect some errors. Each timestamp opens the video at that moment.

Read the full transcript · 3,620 words

0:02 Cool, thank you everyone. My name is Karl, I’m from Autodesk, I work there in the research team working on machine learning in the AI lab, and today I’m going to give you a little bit of a rundown of the things that have been happening with B generation. So I’m talking about generative models that are trained on, yeah, bips.

0:25 So, to set the scene, over the last cple couple years we’ve seen these sort of massive successes with text and image generation, right? So if we ask chat GP, GPT, to write a Hau about machine learning in 3D, it can do a pretty good job, or if we want to, you know, use something like control net to generate images from a prompt and a sketch, it actually can do, you know, really impressive things. U, and I think this is largely, or one of the main reasons for this is because there’s been these massive massive data sets of text and images scraped from the internet, but what we really haven’t seen yet is that happened in 3D, right, because that doesn’t exist, and 3D is still hard.

1:08 This is an example of a paper our lab put out called Make a Shape earlier in the year, and this is this simple kind of canonical problem in machine learning of going from a a flat image to a 3D shape, and you can kind of see, right, like some of the results are kind of not quite sashimi grade, so, you know, you know, even our results, right, like we’re sort of hallucinating a Nick TI on the on the android, and like the the things are floating away, the antenni. And what’s even more challenging is that manufacturable 3D is even hotter, right, because we’re dealing with precise surfaces, things that need to fit together, tolerances, and and the whole deal.

1:47 And kind of the foundation of all of this stuff, and I think it’s been mentioned a few times, is the boundary representation of solid models. It’s the the 99.9%, I think, of all the things we do in CAD. So I want to talk a little bit about the work we’ve been doing over the last, actually, like five years or so, on getting to the point where we might have a generative model that can generate these types of solid models, and I want to get a little bit technical here and kind of run you through, in a in a very delicate Gentle Way, what a B is, and then sort of talk a little bit about how we’re doing generation, you know, using machine learning.

2:24 So B is basically a sort of sewn together bunch of faces, and it all starts by having a series of vertices, like this one shown in pink. You connect two vertices, in this case they make an edge, you connect some edges, they make a wire, you connect, you know, the wire, the inside of that is the face, right? It’s, it’s actually very beautiful. And what’s kind of weird about bs is there are a bunch of different curve types and a bunch of different surface types, and you think about machine learning, usually it’s just like characters, like text or pixels, but here we’re lucky enough, or unfortunate enough, to have many different curve types and many different surface types, and they all have sort of different ways of being formulated.

3:09 Another beautiful thing about the B is they have this really graceful sort of topology of how the different faces are connected together. So this is, on the right hand side, this is what’s called a face adjacency graph, so all of these faces are joined together in this really well-defined graph, and that’s actually helpful, I’ll tell you why in a couple slides.

3:27 So up until about 2019, like most of the work we did with 3D representations for machine learning was actually, we just sample points, with just sample points on a on a B, or we even just render an image, and then we just sort of use some of the existing machine learning models out there to to deal with just sort of avoid the problem, right, just deal with the representation that’s well known. Since around about 2019, almost all of our work is either looking at a graph representation, like the face of graph, or some kind of sequence, right, so we sort of find a way to mash the B representation into like a token sequence.

4:07 And you guessed it, graphs we can use graph neuron networks, which are really really good for sort of learning a representation, and we can use Transformers for the sequences, right, and Transformers are really good for generating stuff. So all of the the stuff you see with large language models, it’s almost always using a Transformer, but graph neuron networks can still be useful if you’re, you’re not doing a kind of generation task, if you’re trying to classify, you know, a cat versus a dog or whatever the the equivalent is.

4:36 And for be up generation, you can kind of break all of the work into two categories, so, and you can think about this is either learning how something was made, so like the steps somebody did, like the buttons they pressed, or learning what was made, so the shape they actually made at the end, and I call these modeling sequence generation versus direct shape generation, and I’m going to talk about these today. There’s a bunch of work that’s happening in the space, but I’m actually only going to talk about our work, because, yeah, that’s, it’s great.

5:05 So let’s talk about modeling sequence generation first. So the first thing we realize is, like, we actually, when you sort of have the sequential data, you can kind of play back the timeline, this is like sort of like a parametric history, but we didn’t really have a data set, you know, back in whatever, 2018, and so that was kind of the first thing we had to do, was actually create this data set, and this is called a Fusion 360 Gallery data set, and it was the first data set at the time that actually released the sort of final B together with the the steps that a user would take to to create that data.

5:38 And it contains basically a bunch of of shapes that are made in Fusion 360, along with that kind of parametric history, and from that point on there was a bunch of work we did, a bunch of work others did, some work looking at how you could kind of learn that sequence, so learn basically like how people would sketch, how they would extrude things. So we had a series of papers, I’m going to walk you through one of those, it has a kind of difficult name, sketch gen, and this, first of all, I, I’ll talk a little bit about the representation that that uses.

6:11 So imagine you just have a simple circle, and you want to sort of feed this to a new network. There’s a bunch of different ways you can do this, the way we do it in this paper is we separate out the topology from the geometry, so we call the topology here, basically, the the type of curve it is, so it’s a line, an O, or a circle, so we have this T in the box, is, is a, is a topology token, and then we have samples of points on that circle that we feed in, these are just XY points that we feed into the to the new network.

6:39 And I’m sort of showing you this not because it’s really exciting to look at how you turn a a circle into points, but we do this sort of again and again and again for curves, for loops, for faces, for extrusions, and you can kind of get the picture that all these things turn into basically a series of of tokens, and that can then be basically fit to this, these sort of Transformer networks. So that’s the process of turning a B into into a series of tokens.

7:11 I’m going to not dive too deep into this, but this is of sketched in architecture, but I think the the key takeaway here is for the first time we’re using this idea of code books, which was used a lot with some of the very very sort of popular, like, DI image generation models, and what we kind of send into the network is a series of tokens, and what we get out at the bottom, that confusing looking line called the cad construction sequence, you can just think of that as the parametric history, right, like the thing you, you click in the timeline of your CAD tool.

7:43 So kind of what we’re doing is giving you, in the end, a sort of editable design that you could go back and sort of edit the, you know, the size of that circle or things like that, and what we found is this can create pretty good results, right, in terms of like generating the 2D sketches or generating the 3D forms.

8:00 And so there’s still a bunch of work going on here right now, like the current State ofth art really just, just looks at sketching and extruding, but just think about all the other operations that, like, you need fillets, you need like shampers, you need revolves, this is just a lot of stuff that that could happen in this space. And the nice thing about this approach is that you end up with this editable parametric model, and so that, you know, for a lot of people, parametric models are a thing.

So direct shape generation, so this is the other approach, right? So here we’re not looking at like what people actually sort of clicked in the interface, we’re looking at the the final shape you get, and think about this like a dumb step file, or IIs file, right, where you have none of that sort of information about how the shape was made, it’s just, just the the literal boundary.

8:45 And a couple years ago we, we published a paper, this is called solid gen, and this was sort of looking at, really it was the first generative model that could actually generate these sort of, directly generate these B shapes. And so again, it all starts with a bunch of vertices, and what Solen is essentially three different models, and it’s essentially like a sewing model, it like kind of sews together the edge model basically picks vertices to sew together, and then the face model picks edges to sew together.

9:17 And so if you have this hierarchy and you build those up one by one, and what you end up with is this sort of sewn together solid model. If we sort of walk through that by looking at the architecture, you can see that along the top here it first predicts a series of points, does that sequentially, then basically the the next model, The Edge model, is saying, hey, I want to reference those points, join up to make edges, and sequentially you do that, and you have like kind of a wireframe, and then the final process is like saying the similar thing, like, let’s point to those edges to say, hey, these are going to be different faces that make up the solid model.

9:56 And the the cool thing is, is when you actually look at this, you look at like how the model is generating these things one by one, it’s sequentially like predicting these these edges and then predicting the faces, and it makes a pretty animation. So what’s interesting about this model is that it’s, at the same time, it’s sort of using these analytic surfaces, right, so it’s using the the cylinders and planes, but it can’t do the more sort of free form stuff you might want to do with, say, like industrial design applications.

10:30 So the the work I want to talk about that’s new, as of this year, and we presented this at sigraph, is a work called B genan, and this is actually, a lot of this is in collaboration with the student Sam, who’s now actually working for us, and this is the first B diffusion model.

10:46 So you’ve probably heard of Mid Journey, or stable diffusion, or do, and these are all, you know, image-based diffusion models, and they’ really really taken like the image gener quality from something which is pretty rough, back in the sort of early days of Gans, to something that, like, you know, that’s actually pretty impressive now.

11:08 And so image diffusion models, if you’re not familiar, they’re this pretty simple idea that you can sort of sequentially, like, introd, like introduce noise to an image, so you take this little kitty cat on the left, you mess it up by adding noise, and what you’re asking the no network to do is basically sequentially sort of learn how to remove that noise, and the intuition here is that’s much easier to do in steps rather than, you know, all at once, and that has some implications, because, you know, the inference time is slower generally, they take a little bit of time, if you’ve ever done Mid Journey you can kind of probably seen this, this sort of diffusion process as it happens.

11:49 And for BS it looks much the same, so we have a couple of levels of diffusion, so one is at the the sort of face level, so we sample points on the surfaces of these these these B surfaces, and we sort of mess them up, and then try and learn how to remove that noise to recover the underlying shape. And likewise for the edges, we, we do the same, so we sort of basically sample points on the edges, mess it up, and then try and remove that noise, and at the end we can then Stitch that together to create these solid models.

12:21 So I’m going to dive a little bit deeper into this, because this is kind of new and kind of fun, and it gives a little bit of a sense of the things we have to do to to get things to work in a neur network. So you may recall, yeah, B has a bunch of faces and edges, and they’re all different types, and the way we can kind of make it a uniform representation is by sampling points. So if you look at the points in 3D, they, they look like this, but if you actually kind of look at them in in 2D, it’s really just looks like an image, right, like so if it looks like an image, we can basically pass it into the same types of neur networks that have been very, you know, well established in the image domain.

13:02 So it’s either a 1D or 2D image, so we have those, and we’re basically squashing these points into sort of 2D bins, and then we run them through, like, these kind of classical variational Auto encoders to get the surface embedding, which is just a signature of that particular surface or a curve.

13:24 And one of the challenges we have to address with BPS is, unlike images, we have sort of fixed, you know, 100 100 by 100, whatever size, B, BBS, we’re lucky enough that they come in all different varieties, right? So we have some really complex models with a lot of faces, some are really simple, and so we need to be able to handle all of those cases. So this particular model has seven faces, and but then each of the each of the faces have, you know, different numbers of edges, and we need a way to handle that.

13:53 And the way we handle that is actually by adding duplicate Edge curves, so in this case each face gets a copy of its surrounding edges, and then likewise for The Edge, it gets a copy of its its surrounding sat and end vertices. And the easy, actually the easy way to look at this is to kind of look at it as a tree, so if you think of this like the overlying structure where these little colored nodes represent the different faces, we’re essentially sort of breaking those apart and giving them their own little duplicate version, and we do that for the edges as well, and so we sort of make it so that it’s a sort of consistent shape of the tree.

14:37 Then we also have to have a duplication to pad out, so neuron networks typically want to fix size, so we have to typically, yeah, typically, have to sort of pad things out, and we do that by sort of adding or duplicating some of the faces, as we, before you sort of pass it to the the network. So that’s sort of some of the dirty little secrets, right, in terms of getting stuff to work, getting, being able to pass the stuff into an, or network, and then once we have that, we can run this through our sort of latent diffusion model, right?

15:11 And so what you’re looking at is the animation of denoising those surfaces, and likewise here you’re denoising the edges of this lamp, and then through a postprocess we can then, you know, rebuild that solid model. So let’s take a look at some of the some of the results, right? So on the left hand side columns you’re seeing basically the denoising process of the the the surfaces, and then on the the right hand columns you’re seeing the denoising process of the of the edges, and this is basically trained on like a Furniture data set, so the the things that are spinning, the Reconstruction of that solid model.

15:55 And so all we’re asking the no network to do here is to just basically give me a, you know, a random shape, and this is the type of thing it creates. So here’s another example, so again you can kind of see this denoising process happening, it’s really fun to look at. And we can train it also on mechanical CAD data, and you get, you know, the types of pots that you typically kind of make, these sort of pris, Prismatic shapes, mesmerizing.

16:34 All right, so, so one of the things we can do is, if we have labels, right, like we have labels for different types of furniture, we can say, hey, give me a random lamp, give me a random bed, and that’s sort of generation based on the cloth. We can also do things like shape completion, so if we have a partial, like, set of surfaces, we can say, hey, give me a bunch of versions of the completions of that that might, you know, fit in with those input shapes, and you can see it’s some interesting stuff happens, right? Like when you give it four legs, it can come up with either a bed or or a couch, and so that’s kind of interesting, right, if you think of kind of autocomplete like workflows.

So I’m going to wrap up and just, like, on a final note, I mean, just like the professor, for me, I’m going to talk about some things that are kind of unsolved problems, just to assure everyone the AI is not going to take over everything. I think the first thing here is like, like it’s not all good, right, like we have some examples in our in our paper, it’s like missing faces, self intersection, wobbly surfaces, and we’re not talking about STL files, this is, this is B land.

17:48 So, first of all, yeah, there’s, there’s issues with Generation, and also just with part complexity, right, like generating complex Parts is a challenge. Also just none of these pots, right now, they’re not aware of any kind of assembly that they exist in, so I think there’s a bunch of sort of Generation stuff that could happen. It’s also been really exciting to me to see, like, some of the discussion around physics, I think physics guided generation is, is, is, is just an area that’s been under explored.

18:14 I think there’s many different ways of going about that, but it’s a really exciting area, because I just think it makes, you know, a bunch of things more kind of robust, you know, existing in the real world. Let’s just say, in the final area, I think, is new interfaces for controlling generation, right? Like, I’m pretty sure none of us want to type in, type in, in text and do CAD, right, like that’s kind of not really going to be super helpful.

18:39 But I, I think there are other modalities, like, you know, sketching things that, like, in this is a paper we had called Sketcher shape, where we did the sort of sketch to to to B Generation. So I think there’s somewhere in between, there is, is a sweet spot, and a way of like kind of using some of the the tools we have today.

18:59 So that’s it for me, this is my LinkedIn, if you want to connect, you can scan, blow up my phone, but yeah, happy to have a chat with anyone after the after the final talk.

More from CDFAM NYC 2024

From Text to Spaceship: Advancing AI in Aerospace

From Text to Spaceship: Advancing AI in Aerospace

Ryan McClelland · NASA Goddard Space Flight Center

Generative Design From Lamps to Lungs

Generative Design From Lamps to Lungs

Jessica Rosenkrantz; Jesse Louis-Rosenberg · Nervous System

Emerging Technology within the Design Process

Emerging Technology within the Design Process

Jenna Fizel; Zoey Zhu · IDEO

Design at All Scales Through Computational Craftsmanship

Design at All Scales Through Computational Craftsmanship

Arthur Azoulai; Diego Taccioli · Slicelab

Modernising Engineering Design Processes with Computational Tools

Modernising Engineering Design Processes with Computational Tools

Dauphin Flores; Sean Turner · Henderson Engineers

A Journey to Digital Prosthetics

A Journey to Digital Prosthetics

Brent Wright · LifeNabled / Advanced 3D

Rethinking DfAM: Across the Production Floor

Rethinking DfAM: Across the Production Floor

Ankush Venkatesh · Glidewell Dental Laboratories

3MF Volumetric + Implicit File Format for 3D Printing

3MF Volumetric + Implicit File Format for 3D Printing

Jan Orend · 3MF Consortium / EOS GmbH

Simulation-Driven Continuous Engineering

Simulation-Driven Continuous Engineering

Neel Kumar · Intact Solutions

Computational Design for Large Gas Turbine Engines

Computational Design for Large Gas Turbine Engines

Bradley Rothenberg; Andrew Kappers · Siemens Energy; nTop

Spherene Metamaterial in Simulation-Based DFAM

Spherene Metamaterial in Simulation-Based DFAM

Christian Waldvogel · Spherene

Additive Manufacturing of Ceramics: How Far Can You Go Using Computational Design?

Additive Manufacturing of Ceramics: How Far Can You Go Using Computational Design?

Alberto Ortona · SUPSI – Hybrid Materials Laboratory

Accelerating Time to Market for Purpose-Built AM Software

Accelerating Time to Market for Purpose-Built AM Software

Marek Moffett; Daniel Hambleton · General Lattice; Metafold 3D

New Advancements in Physics-Driven Design

New Advancements in Physics-Driven Design

Marco Pietropaoli · ToffeeX

Register for Updates and Discounts on CDFAM events.