CDFAM CD/DC 26 · Washington DC · 15 July 2026

Agentic Engineering: Generative AI in structural applications

Abstract

Over the past decade, CORE studio at Thornton Tomasetti has developed a vast array of classical Machine Learning tools for structural design and analysis, which have proven incredibly helpful for rapid iteration in the early stages of a project. The very same ML tools that help our engineers work more productively are now being utilized by agentic systems that orchestrate and execute complex, multi-stage design workflows. This presentation will explore in depth CORE studio’s recent experiments in integrating agentic AI at an enterprise level, along with the safety, security, and compliance concerns such integration entails. We will also examine our work in multi-agent collaboration using the A2A protocol, as well as my personal efforts at bridging the gap between LLMs and CAD software using MCP. These topics are all part of a larger discussion about the evolving relationship between engineers and their tools, and the shifting role of a designer in the early days of the Intelligence Age.

Transcript

From YouTube’s automatic captions, lightly cleaned; expect some errors. Each timestamp opens the video at that moment.

Read the full transcript · 3,167 words

Perfect. Right everybody. My name is Sergey Pigach. I’m a senior associate applications engineer at CORE studio Thornton Tomasetti. This is my second time presenting at CDFAM. I’m very excited to be invited back. And today I want to talk to you about agents, and specifically CORE studio’s latest research into the use of agentic AI for structural engineering applications and CAD. As a disclaimer, everything you’re going to see today still remains very much in the domain of experimentation and R&D.

0:59 We’re not quite ready to let our agents loose on real projects, and there’s a multitude of reasons why though that remains ultimately the goal. And also I’ll share some personal relevance work just to kind of paint the full picture. Now, if you don’t know who we are, Thornton Tomasetti is a global engineering and scientific consulting firm. We have offices all all around the world. We employ about 2200 people.

1:30 And within Thornton Tomasetti, or TT for short, there’s a dedicated group called CORE studio, which is our research and development arm. We’re about 40 people split into different verticals. We do applications development, advanced modeling, BIM, and within that group, there’s also a specialized team called Core AI. And the task of Core AI is to research and develop practical applications of machine learning and artificial intelligence to benefit Thornton Tomasetti and the AC community at large.

2:05 So, at CTBUH in New York last year, I showed a library of machine learning tools that KORE has created for our engineers. And here are just a few examples. So, these are all system level and element level design and analysis apps that really enable fast iteration in the early stages of a project. And all these tools are running on ShapeDiver. There’s a Grasshopper definition that runs in the cloud.

2:31 It handles all of the geometry, all of the business logic. It also talks to our machine learning back end. And because all of these engineering apps are implemented as cloud-based services and APIs, converting them to tools that can that we can just give to an agent proved rather trivial. Especially when MCP came out. And for those of you who might not know, MCP stands for Model Context Protocol.

2:58 It’s just a standardized way for agents to interact with tools. So, when MCP came out, we realized that KORE Studio is actually pretty well positioned to take the tools that we spent years developing for our engineers and just hand them to an AI agent and just kind of see what happens. So, here are some of our agentic experiments. First of all, here’s Bender, named after the Futurama character.

3:28 And Bender’s an agentic system that runs in our AWS back end, but we interact with it through Slack. And you can ask it questions like, “Hey, I’m working in a five-story residential building in New York City. Please help me design a concrete column stack with a concrete footing. Make the rest of the assumptions yourself. Render the results.” And Bender’s going to go ahead and make a bunch of tool calls to very same tools that we give to our engineers to work on real projects.

3:57 And once it gets the results, it can render a little visualization of the column it just designed. It gives you a little design summary. And we can just continue chatting and you know, ask for other materials and other options. Now, the reason why we’re doing this on Slack is because we find there’s to be something very interesting about having agents inhabit shared spaces where people do work.

4:25 So, you know, unlike the experience that we are all used to when you just sit down in front of a chatbot or an agent and you have a one-to-one interaction, imagine a bunch of engineers on Slack in a thread talking through a specific problem, trying to find a solution, and at any point any of them can just at-mention Bender and say, “Hey, can you just run this calc from your old calc?

4:49 Or can you look something up? Or what do you think about this?” And because Bender gets access to the entire conversation thread with all of the images and files kind of embedded into it, it doesn’t need to be caught up. Instantly knows what it is that you’re talking about. And also Bender has access to a bunch of sub-agents that have specialized tools. So, here we’re asking, “Hey, which one of these options meaning steel, timber, concrete has the least amount of embodied carbon?” And so, Bender will reach out to the sustainability sub-agents, which we’re calling Groot.

5:28 And Groot using its tools will tell us that, unsurprisingly, a timber column will be the way to go. Now, we also connected Bender to Rhino Compute, which allows us to visualize geometry by utilizing kind of a series of pre-canned parametric definitions that it can just decide to solve and populate with some parameters. And also Bender can return files. So, it can generate Rhino CAD file. It can return CSV tables, spreadsheets, markdown, what have you and just spit them back in the directly in the thread.

6:11 Now, we’re also very interested in multi-agent interaction. So, one of the emerging protocols that tries to tackle this problem is called A2A, which stands for agent to agent. And the point of A2A is to make sure that two agents that have never talked to each other before, that are meeting each other for the first time, are actually able to communicate and delegate tasks to each other and collaborate.

6:39 And a lot of people also get confused about A2A versus MCP. So, MCP is a standard for agents to talk to tools and A2A is a standard for agents to talk to each other. So, these are not competing standards. These are very much complementary. And the ultimate goal of A2A is kind of to enable what people started terming the agent economy. Right? When you have an ocean of different agents advertising their services, being able to delegate tasks, maybe even being able to reimburse each other for doing useful work.

7:16 So, of course, we did experiments with creating some proof of concept A2A agents, mostly kind of just to familiarize ourselves with the protocol. We didn’t expect them to do anything particularly impressive. So, here’s our POC for that we called Cliff the Surveyor for site surveys. And the reason why we built this was mostly just to probe the ecosystem. Like we wanted to know, okay, so if we take this thing and we put it out out in the public internet and it’s free and it doesn’t require any authentication, like, will people find it?

7:51 Will anyone be able to talk to it? Like, will it be indexed by search engines? And it turned out that no. No one will find it. No one will care. And that’s because A2A is a communication protocol. So, similarly how HTTP tells you how to talk to a website, but it doesn’t tell you like which websites are out there, A2A doesn’t have a built-in discovery layer. So, this is this this brings us to Waggle, which is where I decided to take matters into my own hands.

8:28 And it’s called Waggle because of the little waggle dance that bees do to communicate. So, to be clear, Waggle was very much inspired by the A2A R&D that I was doing at CourseStudio at the time, but this is very much a personal project. It’s not backed or endorsed by Thorn Technologies in any way. They were just nice enough to let me bring this up in front of you today.

8:53 So, Waggle does a few interesting things. So, first of all, it scours the internet looking for anything with a pulse and a valid A2A card. And it adds everything into a searchable index. So, here Cliff the Surveyor our POC that’s now indexed by Waggle. And you can go and look at any details of any agents, their health signals, their quality signals, trust, the list of skills that the agent is advertising, all the things that you want to know when you talk to random public agents on the internet.

9:32 And also there’s an interface to directly talk to agents. So, here we’re asking Cliff the Surveyor to do a site analysis for 120 Broadway, which is our office in New York, and you know, it’s going to think for a while and spit out the report. Another thing is that Waggle doesn’t actually require you to know which agent you want to talk to. So, you can just ask the question to Waggle directly.

10:00 It has its own A2A agent that has access to the index of a lot of agents and say, “Hey, I need this done.” And it will go and find an agent for you that can do your job, delegate automatically to that agent, and return you the results to you. So, here we’re asking like, “Hey, I need a gyroid.” Sometimes you just need a gyroid. And so, Waggle found an agent that can do this for us, and now we can grab the output, the artifact that this agent produced, and forward it to some other agent that can do something else with it.

10:33 So, here we’re asking a different fabrication-focused agent to slice it for 3D printing. And so, we get our Cura slices and a chunk of G-code that presumably we can print. I haven’t tried, so can’t vouch for that. And finally, I mentioned that, you know, part of the agent economy idea is that agents will be able to reimburse each other for doing useful work. So, there are other protocols running on top of A2A like XPR 2 that allows agents to toss each other a few cents.

11:05 So, here I’m using my browser wallet to authorize a small transaction of 75 cents to talk to an agent that is a paid agent. And that enables agent providers to support work that’s a lot more expensive, right? So, here we’re generating a couch asset from a text prompt. That’s an expensive operation, but now it’s economical and feasible because we actually paid this agent to to do that.

11:39 Now, moving back to CORE studio and Thornton Tomasetti, so as a structural engineering firm, we obviously want to know where the current frontier models exist in terms of their level of capability on tasks that have to deal with our core business. So we decided to put together a structural engineering benchmark, which we called Moment Eval or model metrics on engineering tasks. And it contains 91 different problems that we sourced from our talented engineers.

12:10 The prompt we gave them was like, “Do your worst. Like give us the hardest problems you can think about.” And the vast majority of the problems on this Eval are long-form. I think there’s a total of maybe three or four multiple-choice questions in the entire benchmark, and the rest really requires the model to kind of think through the solution, do all the calculations, and output either a single value or a set of values, and then we compare them to a known good answer within 1% tolerance.

12:43 So it’s pretty strict. And because we want to measure how good the models are at doing engineering and not arithmetic, we gave them a simple harness with two basic tools. So there’s a scientific calculator, and then there’s a Python sandbox with three pre-installed tools for scientific computing. There’s no internet access. They can’t look anything up. They can’t download new packages. There’s no engineering-specific software in this bundle.

13:14 It’s just providing models with a way to express their mathematical solution in the language of executable code, and then deterministically solve code to get a calculated precise answer. So, we’re not We’re taking arithmetic out of the equation. So, we didn’t know what results to expect, but you know, given that these were very hard problems that we asked our engineers to come up with, we were kind of guessing that, you know, maybe we’ll get 40, 50% for the model performance in that kind of top tier of capability.

13:50 So, we spent the money, we ran the eval, and it was fully saturated just right out of the gate. So, Claude Opus 4.7 got 96 95.6% closely followed by Gemini 3.1 Pro and GPT-5.5 Medium at 94.5%. And as you can imagine, when we got these results, we got a little bit concerned. So, we started going back in time and trying to see, you know, just looking at older OpenAI models to try and figure out like, well, clearly we missed something, right?

14:29 Like, models got really good at solving structural engineering problems, and we just completely missed when that happened. And the pattern emerged pretty quickly. So, it turns out that the unlock for doing structural engineering successfully is reasoning. Is the ability to have an internal chain of thought. So, the oldest model that would still run in our harness and be able to call tools was GPT-4 it was released in May of 2024.

14:59 And it achieves a very modest result of 38%, which is kind of what we expected. And then 7 months later, OpenAI comes out with O1, which was the first reasoning model ever developed. And O1 just buries 4 all with a shovel. Like, it just jumped from 38% to 74, and then it was kind of just a straight line going to 95% today. We tested in total 50 different LLMs of different sizes from a bunch of different providers.

15:34 And the pattern’s very clear. Frontier reasoning models are really good at solving complex structural engineering problems. Which in retrospect is actually not that surprising because structural engineering is a verifiable domain. It’s basically applied physics and math, and modern frontier large language models are very good at both. So, do with that information what you will. Now, moving on to things that are slightly less distressing. Some of our agentic AI and CAD CAD experiments.

16:10 So, during my previous presentation at CDFA, I mentioned Swiftlet, which is a plugin for Grasshopper. Grasshopper is a visual programming language that a lot of architects and engineers use for CAD. And this is my personal project, but it’s very actively used by Core AI. Swiftlet lets you make web requests from Grasshopper, and it’s essentially an HTTP client. So, Core AI uses Swiftlet to talk to our machine learning backend from Grasshopper.

16:42 This is how all of our ShapeDiver apps that our engineers are using talk to our ML ops pipeline. And recently, I added a number of MCP components to Swiftlet that let you connect your favorite AI agent directly to your Grasshopper definition. To clarify, it doesn’t let you write code your Grasshopper file. Some people get confused about that. It just lets you use visual programming as kind of a no-code way to put together MCP tools that give your agent access to your 3D model, to its data, let it manipulate geometry, and so forth.

17:19 So, here are a couple of examples from Marine Lummus, who very kindly let me share his work with you all. So, he’s He’s very active on LinkedIn. You should definitely go check out his work. But on the left-hand side, there’s a recording of him using Swiftlet to connect an old parametric bridge definition that he had lying around to Claude. And here, he is just using natural language to write code, write design a bridge.

17:45 And on the right-hand side, he connected Codex and Claude to this parametric definition for this NVRDV inspired model to drive the facade geometry and also have the agents create custom dashboards to visualize data, which I thought was very cool. Now, some other surprising uses of Swiftlet and MCP came out of the AC Tech Hackathon that Gore Studio and HKS hosted in LA this spring. So, I got to work with these wonderful people.

18:18 And the awesome part about this Hackathon was that we got to play with their robot arm. It’s a standard bots AR01 robot that has a REST API. And one of the things that can talk to REST APIs is Swiftlet. So, what we did, we ended up creating a digital twin of this robot in Rhino, which is a CAD software. And then we were able to drive the physical robot arm from Grasshopper using Swiftlet.

18:46 Here you kind of see us manually moving around the plane and have the robot tool orient itself normal to to that plane. And then we threw all caution to the wind and connected the Grasshopper definition to Claude via MCP. Now, I jumped out of planes before. I climb mountains for fun. Given Claude unrestricted access to a 70-lb robot arm was about the sketchiest thing I’ve ever done in my life.

19:20 Now that said, Claude Opus 4.8 actually proved surprisingly capable of operating the arm. It was very slow. It had to think a lot about each move, so it’s not exactly practical. But for those of you who understand inverse and forward kinematics and how math goes how much math goes into rotating every joint to get the robots to move where it needs to go, just being able to say, “Hey Claude, move it over there.” and just have it do that was incredibly incredibly impressive.

19:50 So HKS continued working with this robot after the hackathon, and last I’ve heard, they equipped it with a knife. I really hope that they’re just using the rest API and they didn’t just hand Claude a melee weapon. As far as I understand, they’re using the blade attachment to cut some like really cool-looking parametric acoustic panels, but you know, someone should probably swing by their office and just make sure everybody’s okay.

20:17 So in conclusion, I think that the part of the reason why agentic AI has not caused as much of a disruption in the AEC industry is that a lot of the design, analysis, and documentation software that we deal with daily is so awkward and clunky that it’s barely usable by humans, let alone by agents. I’m not going to take any cheap shots at any particular vendors. You know who you are.

20:47 So, when you give an AI system a proper agentic harness with engineering tools specifically designed to cleanly integrate those computational workflows, you kind of unleash a monster. So, it’s also not a domain expertise problem. It’s really an unhobbling problem. Right? Frontier models as as we demonstrated are actually really skilled and really competent at doing engineering work. So my hope is that, you know, with the right tools and the right guardrails and the right oversight, AI agents can help us build truly extraordinary things. Thank you very much.

Register for Updates and Discounts on CDFAM events.