CDFAM Barcelona 2026 · Barcelona · 8 April 2026
Functional AI for 3D Design Automation — From Path Finding to Generative Modeling for Building Construction
Abstract
Great strides made recently in 3D generative artificial intelligence (GenAI) have been propelled by the rapid scaling of large foundation models and advances in generative models such as diffusion and flow matching. However, current neural generators have predominantly been constructed by optimization against image-space losses. Should appearance be the main criterion for 3D design and content generation? Not really. The 3D world we live in is not only to be observed. Accordingly, the main goal for 3D GenAI should be for the generated 3D entities to be used and interacted with, so as to serve their intended functions, just as in the real world.
We introduce Functional AI to 3D design automation and demonstrate its importance and potential for the built environment. Fundamentally, any constructed building, and all the objects and structures therein, must fulfill the desired functional requirements, from architecture and complex structural layouts down to the placements and intricate interplay between mechanical equipment, heat or water pipes, and electrical conduits. Our technical coverages will encompass functionalization of 3D objects and scenes, agentic AI for path finding, and generative modeling of complex building structures, with the ultimate goal of establishing a foundation model for building data with construction intelligence.
Transcript
From YouTube’s automatic captions, lightly cleaned; expect some errors. Each timestamp opens the video at that moment.
Read the full transcript · 3,699 words
0:15 Okay, so well, I’m good to be I think I’m here in a very special conference. I’ve given many talks mostly at academic conferences. So this is a unique experience because number one, I think the speakers are very diverse. I think maybe the most diverse I’ve seen. And also this is very rare that I come to a conference where I give a talk. Most people, if anyone actually know about me, right?
0:41 So So it sort of reminds me of the very first talk I gave at SIGGRAPH. SIGGRAPH by the way is the top conference in computer graphics. Where I had to kind of make an impression. So I feel kind of pressure to make an impression. I’m Richard Sam. I’m wearing two hats. I’m a professor at Simon Fraser University in Canada. And I’m also VP of AI and R&D at Aquanta.
1:03 I’m going to talk about first my my hat as a professor at SFU. So I’m a computer graphics researcher in my training. By now I work very broadly visual computing where we try to understand, process, and generate visual data, images, videos. And I’m very like a 3D person. I work mostly in 3D generation, 3D generative AI. I also now claim that I work on spatial AI just because it’s a very very hot topic.
1:31 But as you can see from my profile photo, I have this fascination about computer computer computational design and fabrication. I think this is kind of a very befitting to the theme of CDFAM. So what’s that picture about? So this was very early work we did 2014 which we called approximate primitive decomposition. Where you do 3D printing, you have something which has overhang and we decided to decompose it into few parts where each part is pyramidal.
1:58 It’s essentially a height field where you have very little waste when you do layered 3D printing. So, that paper have been have received some some press coverage coverage and LSA people they actually made this small Christmas card because this Christmas tree is like a really good example to show you contrast between printing just upside down as a tree would be and also decompose it into two parts where each part is pyramidal.
2:25 And another work I would really like because this So, let me stop a bit because I didn’t want to play that yet. Is this because I talk about pathfinding in this talk and which I’m going to get to why it’s related to architecture. But before that, I worked on this very interesting fabrication problem where you have this 2D region and you’re going to do a path feeling, right?
2:46 So, the question we’d ask was is it possible to use a single continuous path to fill any 2D region? And that was actually possible using something very fascinating. This is actually called a Fermat spiral. Everyone knew about Fermat’s Last Theorem but nobody actually or very few people actually know there’s actually a Fermat spiral where we used that to be able to use a single continuous path to fill any 2D region as you will.
3:10 So, this work was actually featured by Two Minute Papers about 10 years ago and here I have to show you on the and in terms of design, I’m always always fascinated by these really beautiful 2D or 3D problems. You see calligram. This is all 3D 3D printing by the way. Eulerian wires and then the one on the bottom left is actually one we use a diffusion model for a logo design.
3:37 And we also worked on LEGO design. You kind of like hand drawing some kind of a sketch and you’re going to produce a LEGO Technic design and then the first author of this paper actually created a startup. Now it has over 200 people. And the one on the bottom right is this work. This is very Canadian. I will just show you. It’s actually really beautiful. You see what’s going on here.
4:05 The maple leaf. You open it up. You reverse it. I should become become a beaver. So, we were actually very interested in how was what kind of two shapes you can actually do this. We call it the riot. Reverse inside outside transfer. So, these are kind of stuff we actually I did on design and fabrication. So, that’s kind of a pause on the SFU side of things.
4:35 Of course, I work on other stuff as well, but those just really beautiful things that kind of drives me. Okay, Augmenta. So, Augmenta is a Canadian startup company. And we are in the AI for AEC space. AEC stands for architecture, engineering, and and construction. So, our grand vision is imagine in the future you’re able to just use some prompt texts, images, floor plans, and then interior constraints, and then you click a button, okay?
5:02 And you can actually get a full interior 3D building actually works. So, that’s kind of our grand vision. So, inputs multi model prompts, and then you’re going to get a building designed, okay? Fully designed that actually works. So, that’s kind of the key theme of my talk. So, the company was founded 2018 by the generative design team from Autodesk. And my CEO, Freo come Freo is Francesco Iorio.
5:28 He started the project in Autodesk called Dreamcatcher. Right? So, now if you’re in you know, if you’re in computing, computer vision, AI, you know that many algorithms now on generative design or general modeling are called a dream this, dream that. Okay, there’s dream fusion, dream booth, and and and Dream Gaussian. And this kind of started with Google at 2015. They had this project called Deep Dream. Like a Deep Dream actually was not about general design.
5:56 In any case, Friel actually had this word Dreamcatcher. This is how he named his project way before Google did it. So, he was actually the very first person that I know of used the word dream to to signify this kind of hallucination in the generative model. And I joined Augmenta last year. So, I’m pretty new, August, leading the AI effort. So, why did I decide to work on AI for construction?
6:23 So, there are a few reasons. So, in the kind of economic business perspective in the AI landscape, construction is actually a huge industry. It’s bigger than energy, bigger than healthcare, bigger than entertainment, where you see all these C ads 2.0. The result is actually in entertainment, but construction is actually much bigger. Yet, on the other side, the AI adoption is actually the lowest among these industries. So, that’s kind of make it attractive.
6:50 From research perspective, a lot of challenges, technical challenges, that are mostly applicable to 3D generation, reconstruction, but they’re actually amplified in architecture. So, here I’ll show you that in the past few years, the advances in neural 3D reconstruction, that is from a single image, you can have a 3D model, has advanced significantly to a point where I don’t think there’s actually a lot for me to work on.
7:16 So, this is a result by a model called Trellis 2. So, on the left you have a single image, and on the right you have a 3D model reconstructed. Okay, you can see the detail. It looks really just like it’s like probably the the the best you can hope for. But, if you ask question, does it work? Can you actually live in it? Can you actually manufacture it?
7:36 Can you actually build something out of it? The answer for two all is actually no, because these models, how how good they are, they are built using image space losses. Okay, they try to construct a 3D model with some 3D priors so that the outcome replicates or reproduces the input image. So, it’s really about looking right versus working right. So, for construction, building design, you actually want it to work, right?
8:04 It’s not enough just to for it to look good. So, functional buildings go way beyond just looking all right, having this exterior construct that look plausible. There’s the whole interiors, right? And then you have the the layout, you have the walls, you have the windows and doors, and how they work, and where opens, and how do you access. And also, finally, you have something that is called MEP, which is mechanical, electrical, and plumbing.
8:27 You have to put all of them together for the building to work. So, this is actually much more complex than just having something that looks all right from the outside. So, then, this kind of goes to functional AI because I talk about functionality. So, how important is functional AI? Now, let’s think about 3D gen AI. You ask, why would anyone want to generate something in 3D? It’s not just look at it, right?
8:51 It’s actually to use it. Right? I hope most of you agree. And the world model is something that everyone talks about now. There are many startups, including Yellow Cone, working on world models. So, what is it really? You can ask, you know, ChatGPT, but people mostly agree the world model is about a model that tells you how the world works. Okay? So, in this sense, I believe all of these sort of hot buttons people are pushing these days, especially where embodied AI, physical AI, and world model.
9:19 I think what underpins all of them is actually functional intelligence. How things work, what they do, and how to use them, how the world functions. And then, if you go down in terms of like these fundamental components like research topics, structured understanding, structural representation, physics, interaction, and also motion. So, these are the four to me four pillars to support functional intelligence. And going back to another challenge for construction.
So, to come to in to contrast with manufacturing. So, I want to take again three of my CEO. He he he in in one of his meetings he put up this iPhone. He said, “Okay, if you design one of these, you can make a million copies of it. You can make money.” Right? But, buildings are not meant to be the same. Where really do you want to have two buildings that are exactly the same?
10:13 So, the the the cost of construction does not amortize over the volume of units you make. So, that’s kind of difference between construction and manufacturing. This is especially true with the kind of buildings that Augmenta is interested in. We’re actually not interested in residential buildings. We think it’s kind of more predictable. We’re interested in mission-critical commercial buildings such as schools, hospitals, and data centers. We want to fully engineered constructible buildings designed for these purposes.
10:41 And we’re proud to say that Augmenta actually made history last year. This is a elementary school in Michigan. That was the very first time this this the whole building’s actual system was made by AI. Data scarcity is another challenge. So, there have been some large-scale data out there for architecture, but some of them this one was actually very very new, 5 million, but their LOD allowable allowable detail level is very low.
11:12 They’re basically essentially just like exterior, very rough abstraction. There’s no interior, so this building cannot work. And another data set is synthetic data set. It’s of commercial buildings, very small buildings. They don’t have the kind of complexity that we aim for. And some companies actually decide to scan physical buildings. It’s my years and and a lot of effort to scan physical building, but the problem is these buildings will never be able to be made public.
11:44 Right? Because there’s no way that that you have allowed these building structures. And also you cannot really scan MEP. You cannot open up all the ducts and actually scan the scan the interior. So, their ownership and and data scarcity are are issues. And what’s more, construction is not only a multi model problem, it’s a multi-trade problem. Right? Even for MEP, you guys have three classes of engineers.
12:05 They actually have different expertise. They have to talk to each other. They have a lot of resolution they have to they have to resolve. So, in reality it’s a much more complex problem than just designing like a chair. So, among all of these architectural, you know, structural and MEP, which one is the hardest? Again, I’m asking ChatGPT and and a lot of a lot. A lot. So, it’s they are agreeing, of course, so that’s a that’s a good sign.
12:32 So, actually turns out MEP is the hardest. It’s a harder than architectural, harder than and structural, which is interior layouts. Because it’s most time-consuming, require most coordination. So, that was one of the reasons why actually Augmenta decided to tackle the hardest part, which is the MEP. And so far with our path planning algorithm, we’re able to complete a 16,000 mi of conduits. And Augmenta has the only in market actually we’re selling we’re making revenue.
13:04 The only in market solution that has a full automation for MEP. And your input is building geometry, sources and targets, and some scheduling constraints. And with that, we have the only MEP simulator, 3D simulator that actually be a data engine. So, we have the training data that we own. We can actually train these target models to very quickly produce MEP outputs given the geometry. And our solution is a fully agentic.
13:32 Probably I won’t be able to say too much about it because of confidential confidentiality. I’ll only tell you what it does. And and and it’s also has easy integration with with Revit, which is from Autodesk. And we have also routing guidance, which has some level intelligence. So, in the end, we’re able to reduce work by these MEP engineers from weeks down to a few days, but we still have a lot to go.
14:00 But, MEP is only a start. So, I think the last speaker talking about foundation model, we also want to build a foundation model for 3D buildings and and construction. So, on the data side, we want our buildings to have geometric fidelity, full interiors, not just exteriors like like let’s say the building world, which has only 2D polygons abstractions. Want to have functionality, right? And for that, we need to have MEP, we need to have the layouts, which are all functional.
14:29 They It’s not enough for them just to look all right. I’ll come to that. And finally, multimodal annotation. Want to connect the geometry and the space as the spatial elements of a building with the foundational knowledge from these large language models or large visual language models because they have knowledge. But, the problem is that we need to align the spatial constraints with these LLMs and VLMs. And my approach is actually not to to physically scan them because of ownership issues, right?
15:02 We actually want to synthesize this data set, right? So, with that, we have ownership. And also, controllability is another another key advantage. So, I worked on autonomous autonomous driving before. So, some sort of companies came to me and said, “You know, when we drive this car out, we can very rarely capture down truck, a car carrier, or the accident scene. So, these are very, very important. This out of this distribution data.
15:29 So, the answer is actually to synthesize them, right? So, that’s why controllability is something you can actually obtain using synthesis. What it generates these buildings at 3D out of the level three three plus, and also with the special alignment with between the special and the large language models and visual language models. Also, that will have this functional and constructive intelligence. And just as a glimpse of kind of our approach, I’d like to start with floor plans.
15:57 So, floor plans can be generated using like double banana, right? So, I’ll show you how how good it is. And then we do lifting, and then we have these interesting problems called the multi-floor consistency in contrast to multi-view consistency in 3D vision, and also functional rectification. Just as a teaser, I’ll show you if you were to use a text prompt and ask Nana Banana Pro to generate a floor plan, it looks actually pretty good, right?
16:19 How quickly can you identify something that’s wrong? Okay, so this storage room has no door, right? So, you need to be very, very good actually to to to actually identify this. How about this one? This one is much, much harder, right? I want to generate two floor floor a two-floor floor plans, and then if you are really keen or fast, you can really realize the first floor has two men’s rooms for some reason, right?
16:48 And only one women’s bathroom, right? So, these are things that these, you know, current models will give you. They look all right, but if you you get a close examination, something’s wrong. So, we need a way to actually functionally rectify them and correct them. So, that’s the first step. We we have some idea to do how to do it, but I’m not going to say too much.
17:09 Even before Augmenta and ASF, you I worked on AEC as well. So, I was actually fascinated by programs, right? So, CAD programs or building programs one reason is because we know transformers are very powerful, right? I mean, programs, right? It’s kind of purposely built for transformers because it’s all about next token prediction. And also natural language and space or let’s say visual data has a gap, right?
17:40 So, to fill that gap, we decided to use domain-specific language or DSLs, right? So, I believe that these DSLs are the natural conduits between natural languages and what you want to do in the end. So, the next two projects I’m going to go through very quickly highlight this kind of approach. The first of this was a CVPR 2025 highlight paper where we did develop a domain-specific language for these abstractions for architecture, right?
18:05 And with that DSL, we’re able to develop a transformer model that is able to predict this DSL from very sparse point clouds, right? So, that we’re able to produce the best algorithm out there that’s able to take very sparse point cloud and produce an architectural abstraction as you can see over there. And then we didn’t stop there. So, as a follow-up, we tried to work on more sophisticated higher LOD level architecture.
18:32 So, we developed a novel representation that allows you to do these multimodal input to architecture output. You can use text, sketch, and also single view images. And then this representation we call it VLT or ArcVLT, and VLT stands for visual language twin. So, it’s a twin representation. When people talk about digital twin, it’s usually about, you know, what’s what’s digital and what’s physical reality. So, here we have visual representation, parametric representation, and then we have language representation, which is a DSL, and they have this dual or bidirectional ability.
19:07 So, here’s a short video to show you what happens on the in the middle is the visual and what’s on the right is the program. Then you have whatever you do here program. Users can perform direct 3D geometric edits in the left viewer, which are synchronized to the program parameters on the right. Conversely, users can precisely adjust numeric parameters on the right. The corresponding elements are highlighted in the 3D view and the geometry updates in real-time.
19:35 This demonstrates Arc VLT’s bidirectional synchronization between geometry and parameters, enabling twin space editability. All right. Okay, so I’m I’m done. So, I’m going to just maybe leave a few questions. Again, I’m here. I don’t know any of you actually. So, I’d like to leave some questions. Maybe it will be a food for thought and maybe it will start some kind of a a discussion. I believe the visual language twin idea should be ubiquitous in design.
20:04 Not only for architecture, it could all be in MEP or any other kind of design, right? Again, because we believe that transformers are powerful. They’re good at generating programs. Data is an issue, which I’m happy to discuss. And we decided to use generative design for our functional AI approach. I believe that’s also something that can go beyond construction. And finally, I want to talk about foundational models.
20:34 If you go back to the original Stanford paper, they call it foundational foundation models. Actually, there’s really nothing foundational about it, right? So, I think I’ll call it foundational model because these models are trained by like broad large-scale data and then can be applied to many downstream tasks. Actually, it doesn’t have really foundational understanding of space physics. And this is a access that I do every month now.
20:55 I just pass this image to any of the large foundational models nowadays. I ask how many fingers is that? Five. Every single time. Five. Five. Five. Right? So, just goes through to show you that it’s actually kind of just produce a result it’s expect to see. It’s actually not really looking at it. Right? So, I believe that if you want to build a functional model, you can actually look into actual functional knowledge from the brain.
21:18 So, I don’t have time to actually go over this, but there’s some recent studies from cognition telling that babies were not born with a with a clean slate. They are some pre-programmed structural understanding of the world. Right? And also, there’s work on there’s a special memory in our brain using like place cells, speed cells, grid cells. These are all things that are lacking in these large functional models, VLMs.
21:40 So, we should look into them. That’s it. Thank you very much. To learn more about the CDFA.com Computational Design Symposium, access the archive of previous presentations, interviews with speakers, and information about future events around the world, visit CDFAM.com.
More from CDFAM Barcelona 2026

Artificial Intuition: Building an AI Mind for Electromagnetic Design and Engineering
Mike Frei · ARENA Physica

Design for real world engineering: integrating uncertainty into product assessment
Greg Grigoriadis · Metisec

Large Engineering Models: Reimagining Design, Simulation, and Manufacturing | EMMI AI
Quercus Hernández · Emmi AI

Beyond the Specialist: How Istari + AI Expands Who Gets to Drive Innovation in Hardware
Rebeka Melber · Istari Digital

Digitizing Body-in-White Development with MAS Synera
Juan de Dios Escribano Felguera; Tilman Steininger · SEAT; Synera

HOOPS AI: Correlating CAD Geometry with Manufacturing and Business Process Information
Luis Salazar Betancourt · Tech Soft 3D

SubSimX: Interactive Subdivision-to-FEM for Computational Design
Johannes Müller-Römer · Fraunhofer IGD

Empowering Architects with Early-Stage Environmental Intelligence
Michele Pescatore; Carol Fanjul · AiA Life Designers

Bridging Data to Geometry with Implicit Modeling
Wesley Essink · Siemens Digital Industries Software

An Engineer’s Approach to Integrating Machine Learning in Generative Design Tools
Thomas Rees · ToffeeX

Conjugate Heat Transfer Optimization for Turbine Blade Thermal Performance Using Field-Driven Design
Max Gaedtke; Markus Lempke · nTop; Siemens Energy

NeuralShipper: Generative AI for the Next Generation of Ship Design and Manufacturing
Shahroz Khan · Compute Maritime

Beyond Parametric: Explorable Simulation for Real Design Iterations
Laurence Cook · Generative Engineering

Architected Porosity Informed by Real-World Data for More-Than-Human Thermal Comfort
Maria Claudia Valverde Rojas · University of Stuttgart, IntCDC

Real-time Multi-Physics Collaboration for Real-world Engineering
Nikolas Borrel Jensen; Oliver Littlewood · Pasteur Labs










