CDFAM CD/DC 26 · Washington DC · 15 July 2026
Authoring Autonomy
Abstract
When the core value proposition of the humanoid form is generalized capability through massive retaskability, is the world’s most dynamic humanoid hardware enough? The next industrial revolution will rely on industry-leading AI brains in addition to production-grade humanoids and scaled manufacturing. These brains are taught with large data sets rather than programmed, but spatial reasoning can’t be scraped from the internet like text and language (yet). It demands a novel ecosystem of agentic software sandboxes where users can teach robots through physical demonstration, naturally describe tasks, and interactively refine generative robot behaviors and application interfaces.
Transcript
From YouTube’s automatic captions, lightly cleaned; expect some errors. Each timestamp opens the video at that moment.
Read the full transcript · 6,479 words
0:25 This is Atlas by Boston Dynamics. A generally capable humanoid robot. A doer of things, as we like to say, because in theory, this is a product that can do anything. Well, almost anything because I don’t want to be disingenuous. Our focus right now is on industrial applications. Everything I show, every design decision we’ve made, everything about this presentation is in service to this robot doing real work in a factory environment.
0:54 You got to start somewhere. So, in my circles, you’re either working on a humanoid robot and it’s the only viable way forward for robotics or you’re not and it’s an extremely overengineered solution in search of a problem. And I think I’ll admit that a lot of the kind of gaps that we’re looking to tackle in these factories, those could be solved with the traditional types of industrial automation methods that exist today.
1:18 It just wouldn’t be cost effective to take each and every automation problem in a factory and design a bespoke robotic solution for that. So to be clear, the value of Atlas, this isn’t a technology breakthrough. This is an economics breakthrough. With a single investment in a generalized piece of hardware, you turn all of those factory gap automation problems into a software problem. And if you can turn something into a software problem, it becomes much much cheaper to kind of solve all of these different things.
1:52 So this is this is the fun of, you know, building a machine that can allegedly do anything. And while you, the customer, don’t have to worry about that hardware piece, we still very much do. We still have a two-pronged problem. We have to create the world’s most capable and generalized piece of hardware that can operate in humanpurposed environments doing humanpurpose tasks that’s capable of doing things we don’t even know it needs to do yet.
2:22 And we also have to solve the generalpurpose software problem because you probably don’t want to build a custom application for each and integration right there a lot of layers in the software there for each and every factory problem that you work on. So I want to talk a little bit about the most visible aspect of the robot which is the industrial design and the mechanical design decisions we put into that.
2:46 One of the first things that you know we all had to kind of sit around and ask is does this thing need a head? What does it mean to put a head on a humanoid robot? You know, do we need it functionally? And I think I think it’s a pretty easy argument utilitarian argument to say, well, you need to have your perception, your robot cameras roughly at eye level again, if you’re doing humanpurpose tasks in human purposed environments.
3:08 That makes a lot of sense. But why give it a face? Why allow it to articulate in this direction? Why can it articulate in this direction? What is the point of that? Well, the point of that is we’re putting this machine into environments with you, with people. So, this machine needs to be able to communicate in a way that we’re used to communicate, which is looking at each other’s faces, right?
3:32 So, if you walk into a room and a robot gazes at you, you have a sense that it recognizes your presence and it’s going to behave accordingly, which is safely. If the robot looks at a workpiece on the floor, it lends predictability to its next action. It’s probably going to try to bend over and pick that piece up that perhaps it dropped. So in this way, it’s very intentional and it’s very worthwhile to take the time to put a head on this thing.
4:01 So, we have a head, we have two arms, we have two legs. This is the morphology of a humanoid robot, but it doesn’t look like a person. And it’s certainly not meant to, right? We’re not trying to trick you into thinking this is like you, that this is a person with the type of agency that a person has. What we want you to do is we want you to look at this thing and say that’s a useful tool.
4:27 That’s a piece of industrial automation equipment. And recognize that from the language of the design we’ve used and the behavior of the robot itself. And also I don’t know why we would try to replicate a person because people are limited. If you’re going to build a machine, let’s build a superhuman machine. That’s the fun of doing this. That’s why we’re in robotics is because with the intelligence of human designers and machines, we can achieve more than we could have achieved before.
4:58 So this isn’t our first rodeo. We’ve commercialized a couple of other robotics products. You might be familiar with Spot, the yellow robot dog, that focuses on mostly industrial inspection. I’ll touch on that a little bit later. We’ve got a stretch robot that works in warehouse environments for logistics work. So we know that building the world’s best-in-class hardware isn’t enough to have a successful product. You have to have it has to be extremely cost-ffective.
5:21 It has to be extremely reliable. So how do you design for that? Well, you simplify simplify simplify. You limit parts. You mirror things and you modularize. So there are only two actuators outside of the hands which are kind of special little robots themselves. There are only two actuators used across the entire robot. So the design language becomes very repetitive and blocky because you’re casing the same sized actuators.
5:45 A lot of the limbs look very similar. Like there isn’t much of a distinction between an upper limb and a lower limb. There isn’t much of a distinction between the way an arm is designed versus the way a leg is designed. And part of that is we just want them all to be the same kit of parts to make this thing simple and field replaceable. There isn’t in a lot of cases there isn’t really a notion of a front and a back especially with like the knees and the elbows.
6:08 And you know that’s that’s on purpose. That contin infinitely continuous rotation in those joints allows for things like emergent behavior to arise. Allows for the robot to work in kind of tight factory spaces in ways that are perhaps more optimal than would be possible with a human body. So to get all of that optimization, we have to turn to the software. Great. You know, you’ve built the Ferrari of robots.
6:37 How do I actually teach this thing to do something useful? And this is this is the like this is the big thing in the industry right now is you know VC money pours into building a bunch of humanoids. But you know how many useful things are they capable of actually doing in the world? It’s not an easy problem. But before we actually get into the authoring of Atlas, I kind of want to talk about this journey of getting to this point with robots because when I started, I’m not like a product manager by training or even by experience.
7:05 It just sort of, you know, these things just happen. So when I got into this position and we were about to launch orbit which is our cloud software for managing factory applications and managing Boston dynamics robots. This whole thing started in co actually we were like no one can see our robots. Our sales team needs to be able to demonstrate these robots. What do we do? So we build a remote piece of software so that you know people sitting at home can get on their computer.
7:31 They can drive spot around remotely. And then we’re like, “Oh, this kind of seems like the point of robots is to be able to remotely operate them.” So co kind of accelerated that realization. Then we turned that into industrial software, but it was always through the perspective of the robot. It was this the software was about driving a robot. Well, who cares about the robot when you’re building robotics products?
7:50 It’s a means to an end. It’s supposed to be an invisible background agent. So we decided we need to pivot and we need to make the software about your facility. If the product is managing factories. So with spot, we’re largely in the EM space, the enterprise asset management space. We’re going around and running inspections to maintain uptime. We’ll get into Atlas and Atlas is about doing real work.
8:12 So the big thing was well, okay, how do I make software about the world? And I had started in construction. My very first project with Boston Dynamics was working on an autonomous scanning solution with Trimble, the geospatial company. And we were going out to job sites and we were saying, “What if there was just this machine that constantly roved your site and anytime you needed to see updates in terms of what was actually built, you know, reality versus spec, you could go in and you could verify that.” And we pushed that product.
8:47 And I very I believed very firmly at the time that in order for a robot to do any I wanted to get to robots building things. You know, it wasn’t like my dream to use robots as like an autonomous tripod. I saw that as a stepping stone to robots doing real work. And so we said about recording the world and I thought you have to record the world to do real work.
9:08 Like the robot has to know this thing is right here and it has to be able to move within some precision of it. And I’m thinking very like old school slam at the time whether it’s visual or adometry based and that led to the thinking behind orbit. Also we’ll broach the subject of digital twins. There are probably a lot of people in here that are you know way more qualified than me to talk about this but I’ve always been struck by this book.
9:31 It came out in 1992 which I think is very special because that’s the year that Boston Dynamics spun out of the MIT leg lab. And this was really the first book where a software engineer was like, I think that software is going to get good enough that it will reflect the actual current state of the world. That was a novel idea in 1992. And sure enough, in 1993, Xerox Park released the first map viewer.
9:57 I can query the internet and I can be like, I need a current map of a place. And today, we take all this stuff for granted, right? We have, you know, the compute mobiley, da da da da da, and we’re not querying for a map. We’re actually commanding a robot. We’re saying, I need this robot to convey me, the person, transport me from point A to point B.
10:20 So to me, this this kind of shows in a in a very kind of like fluffy and simplified way the promise of digital twin methodology and software. We reflect the current state of the world and we use that reflection to control the current state of the world. There’s all that fun like simulation talk that can happen in between that gets sandwiched in there where we use software to create a parallel world to kind of test and and predict but we won’t get into that too much and I’m sure some of you probably will anyway so so we you know we presented this interface and went very simple low barrier to entry 2D blueprint upload pin things I’m very proud of the little pill that I made that that’s an abstract kind of representation of spot and you watch the little They’ll walk around and this is how you monitor.
11:05 This is how you reflect the state of your factory and it also allows you to see what’s the current state of the robots, what’s the current state of a gauge reading and alert me to any issues. So in practice the kind of concept of operations that gets set up is you have people centralized in a control room. They’re in control. You have a bunch of robots. These robots just live as infrastructure in your facility.
11:27 This will be true of Atlas as well. They’re just all docked and they’re on the network and they’re waiting for something to command them, whether it’s a scheduler or an agent. They go out and they do whatever inspections need to be done. They report that information back centrally to the people who manage that work. They flag anything they’ve been programmed to flag. And then ultimately you use that information to direct work.
11:48 And you might direct human work or you might kind of close the loop of the twin and say, “I’m going to send a robot out. Maybe I want to reinspect. Maybe it’s 3:00 a.m. And before I send a maintenance person in because no one’s in the factory right now, I’m going to send a robot to like double check. But ideally, you want to have the robot do stuff, right?
12:06 And this was always where we were trying to get with with Spot. Spot has an ARM model. It’s useful in certain contexts like public safety and things like that. But again, all of this is like how do we give robots agency? How do we talk to them and do things? And you know, we when I started working on Atlas, we were electrifying Atlas and Stealth. We knew we wanted to try to productize the humanoid.
12:33 And you know, we’re looking at things like again, we’re setting up, we’re using all the same moves. Like nothing about this software is different from how I would approach, you know, doing something with with Spot. It’s all it’s all very much like workflow based. It’s like, what’s my robot doing? What’s its state? Where is it? Is it doing the right thing? Da da da. But there’s not much flexibility.
12:53 This is all kind of premised upon the robot doing one single thing. In this case, automotive sequencing. We’re doing like part sortation as like an automotive intral logistics task. It’s super tight. So, we have this application scaling problem. It’s like, okay, well, if we’re working on this robot that can do anything and you want to switch the thing that the robot’s doing, you have to go in and and we have to build a completely new suite of software.
13:16 That sounds like a lot of work. It sounds like we haven’t really solved that problem, right? So again, very traditional interfaces. It’s all about it’s all about job tracking. So to answer this problem about how do you generalize software, you kind of have to look at the history of user personas of software. Specifically, the personas of the creator of a piece of software versus the consumer of a piece of software or the author versus the end user.
13:41 So traditionally the author has been in an IDE, a specialized piece of software for creating software largely textbased in some cases you get a little mini preview of what you’re doing. Over the last like I don’t know like 20 years or probably honest honestly since IDE existed people have been experimenting with low code or no code. And what that does is it lowers the barrier to entry.
14:07 So it starts to blur the line blur the line between the author and the end user. But more importantly is it starts to put authorship in the context of your application. So it starts to speak more to extensibility and modifying your existing environment than kind of creating from something from scratch. You may have heard that you can talk to your software now. So you know this has completely changed everything.
And in my opinion, there’s there’s really no longer any meaningful distinction between an author and an end user. And most importantly, all of a sudden, any software environment becomes capable of doing anything, which is a very confusing premise. It really feels like completely blue ocean in terms of what software is, what’s its cap, what it’s capable of, and what it means to kind of engage with it in terms of that persona level.
15:02 So in theory if you can talk has anyone by the way has anybody used like Figma make or claude design has anyone kind of yeah or just use claude to make a website right or to you know make Dwan’s slide that was missing a piece cuz you know AI you know it’s not perfect yet but you know okay well if you can do it with software surely you can do it with machines how hard can that be well let’s find out the first step is does the robot stand?
15:33 If you make a humanoid, people kind of expect that it can stand up and balance, but that is a problem in and of itself. So, we’ve moved from something called model model predictive control to RLbased control. So, essentially 100 times per second. Well, before we get there, you start with some kind of goal. Perhaps it’s a backflip and you’ve done motion tracking or an animation. You’ve provided some kind of example and you say this is the perfect motion or it’s natural walking which is a thing we really care about and you say this is the example I want you to be able to do this perfectly I’m going to simulate this you know a bajillion times and then 100 times per second when I run this policy you’re going to observe the world you’re going to observe the state of your actuators and your relative position and you’re just going to kind of like make the next predicted move and it has the illusion of natural walking or a backflip and if you’re if you’re working on this technology, you you know, you have to have a little bit of fun with it.
16:29 We’re very proud of doing our acrobatics, even though we get we get a lot of crap online for having fun. You can do both. You can work on useful applications and make your robots do back flips. So that’s great. Your robot can stand up. That seems like the minimum product expectation you might have, but it needs to do stuff. And a lot of the hard problems in robotics today center around manipulation.
16:52 So, if we have a whole body control system with RL, that doesn’t really solve all of the fine manual dexterity. It might someday. We could get into this whole thing about, you know, everything I’m saying whatever, right? A year from now, this technology will change and we’ll figure it out. But this is what’s working best as a stack right now. So, in this case, what we’re doing is we need to be able to demonstrate to the robot the exact kind of manipulation behaviors we want.
17:21 So this is called learning from demonstration or imitation learning or behavior cloning is what I use most often. And what you do is in the case of like automotive sequencing here you suit up in VR you demonstrate all of these things to the robot and if you do that a bunch of times then the robot approximately 30 times per second looks at the state of its motors observes the state of the world and takes the next predicted motion.
17:44 So we call this pixels to torque which is a kind of a funny concept which is just that the main input there is you see some pixels and then the predicted output is your motors move in the most statistically likely way relative to that scene. So it gives the illusion that the robot knows what it’s doing and understands its surroundings. But if it’s not conditioned by anything then it doesn’t.
18:08 It’s just this kind of statistically corresponding motion based on what the robot’s seeing in its environment. But it works. You demonstrate a thing enough times, the robot will do it. And you have to have more and more data to generalize. If if I demonstrate, you know, picking up my phone off the podium, great. If I say, you know, now my phone’s on the chair and I run that same policy, the robot’s not going to know what to do, right?
18:32 So, it’s very limited. And that’s a problem for us, right? We need this thing to be conditioned with reasoning. So, we have to go up another level to that reasoning model. And we have to start to look at the world of VLMs. And what VLMs allow us to do is say we have a representation of an environment. The way we want to be able to talk to robots is the same way we want to or the same we want to be able to talk to them as if we’re training employees.
18:57 So we need to be able to say like everything that’s labeled in here. An outbound supplier parts container, the parts that it contains, the fact that there’s an exception behavior. If you drop a part, you need to take that out of production and here’s what to do. Telling it to do something three times versus doing it zero times. All of these things need to be communicated through language.
19:18 So, we achieve this through visual language models. These are models that can kind of understand the correspondence between language and objects in the world. That allows us to give language commands to robots or really the agents that control the robots. Those agents can then understand what we’re talking about and through a VLA conditioned behavior cloning model actually affect that change. So that ends up being the stack RLBC VLM for now.
19:43 There’s worlds in which you know any number of those things could converge. I would I would in fact bet on a lot of convergence in the future between these methods but again this is what gets results now and allows us to push the product. So you get this thing and we’re like okay we’ve done a behavior cloning model for automotive sequencing and you’re like well that’s that’s great.
20:05 Maybe I’m not in automotive and I want to do something else. So how do I do that? How do you augment this thing? And this is this is again we don’t have the advantage of an internet scale corpus of text to train a a large language model. We need action data which has to do with proprioception or which has to do with perception which in some cases has to do with tactility and force feedback.
20:25 So that stuff’s not just out there. We have to produce that. And so much of the industry right now is at that stage. How do we get a large enough amount of data essentially about the action space that we can start to generalize it using the same methods that large language models use? So the way we’re actively doing this right now is by switching interfaces from web- based interfaces to VR.
20:51 Essentially, we need to physically embody the robot and use our bodies to control the robot’s motion. And we need to scale that up really fast. Like I’m talking like giant factories full of people decked out in VR doing a bunch of tasks and an efficient data pipeline to collect all of that. And generally the way this works is you suit up, you have a whole body tracking system.
21:14 You have your controllers of choice. You slide on your headset and then you’ve got pass through. You see like I am here, the robot’s there. And then you launch VR and now the VR photosphere is the robot’s perception. Now you are the robot. Now, any motion you take will move the robot. That creates some funny situations. You can’t really like be touching buttons that have to do with the UI anymore.
21:37 Now, your whole body is consumed by what you’re doing. If you’re using controllers, you can move your fingers around because we can abstract the grasping. If you’re using gloves and you want full control, I don’t know. We’re working on that. That’s a hard problem. We You start to get like really weird ideas. You’re like, “well, the robot the robot doesn’t have a tongue, so maybe we could put like a button in your mouth and you could like touch it with your, you know, the robot doesn’t have a tongue yet.” no, just kidding.
22:05 We’re not doing that. We’re we’re 100% not doing that. This is recorded. So what does this look like to suit up? You know, again, you’ve got the headset, you’ve got typically controllers, so you can use the buttons to grasp and the and the buttons themselves tell you where the controllers are. You’ve got a chest tracker. You’ve got faux trackers. And then you’ve got, in this case, beacons.
22:26 And there are a couple ways to do this. Like this is an outside in tracking system because you need line of sight to see all those trackers. There’s a whole emerging set of technologies that use IMUs and even magnetism so that you can run a server on the headset itself and you can track the relative location of those trackers without them being in line of sight. So, that’s good.
22:46 And it’s especially good, too, because sometimes you need to reach like this. You can’t keep your whole body in view of the VR glasses even if you had an extremely wide field of view. So, you know, chipping away at problems kind of embodiment problems there in terms of how you control all that stuff. And this what it this is what it looks like. This is this is not me for for what it’s worth, but I did consider wearing this outfit for this presentation.
23:12 There’s a lot of squatting involved, so and it gets like it’s like it’s pretty hard work. So a lot of these shots it’s like like shorts and like tank tops and people are like sweating. It’s it’s very physically demanding work over time to do it. And the goal is to make the system so that you perform the task naturally. If you have any kind of consciousness, any cognitive overhead of the fact that that the robot moves a certain way and you have to be careful about how it moves or the controllers are slightly awkward and you kind of have to make up for that gap.
23:43 That’s worse data. And you know, someday we’ll have enough data that bad data won’t matter, but for now at the scale we’re at, data quality is, you know, is is very important. So, yeah, you can see this is actually us in a Hyundai plant doing real sequencing. So, those are roof rack parts for a model. And you can see the the demonstrator there kind of going through.
24:09 So you do that enough times and it works. And it’s like I said, it’s a part of part of this is we have to we have to train the people who are training the robots. We call them pilots. And there’s a little bit of latency. We’re working on all that network problem. So there’s latency when you put in when you move your hand to a part versus the robot moving its hand apart.
24:29 It doesn’t take much to to notice that. So we actually do a lot of the training in SIM because there’s no latency in SIM. So you can run this thing in whatever, right? You put on the VR headset and essentially you’re demonstrating and you’re selecting a host. That host could be a physical robot. That host could be a PC that that runs some kind of sim environment.
24:51 SIM has a little bit more overhead cuz you have to have all of these, you know, all these special meshes made to to accurately represent the environment. But we can also create models from simulation data. And we can also co-rain. You can have a mixture of simulation data and data from the robot’s hardware. So how do you actually you know how do you improve this right? So we’re kind of saying like all right you have in theory you have some kind of base behavior model you can add data to that and retrain that but when you add data to something just kind of like dumping more data in it there’s some problems with that that’s not the best way to work so there’s a method called dagger that’s really popular with this type of teley operation and what it essentially means is that you’re a supervisor in the loop you run the policy the robot begins to do the work autonomously and if you see that the robot’s about to make a bad grasp or you see the robots moving in some way that’s not optimal, you can pause it and you can take control and then you can complete that move and you can release control back to the policy.
25:54 And by doing that, you’re annotating, you’re saying like there’s something out of distribution right here between these two action chunks with this timestamp. So you give context to the behavior that’s out of distribution and that is a much much better way of essentially retraining these models. And then you have this positive feedback loop where you do this over and over and over again for each new task that you demonstrate.
26:18 And that’s how you build up that data set. And part of the principle behind some of this stuff is, you know, is a customer going to be running Dagger in their factory? I mean, hopefully not. It’s not particularly userfriendly, but it’s really hard to say ultimately how this stuff will manifest. And especially in the early days, yeah, we probably will be on site, you know, collecting data with the customer.
26:41 But you can imagine that if you’re a customer in the future and you’ve invested in this robot and you’ve kind of gone through some process of describing a new task for it to do, you’re going to want to be able to test that, right? So maybe it’s a full SIM stack. I think that’s the ideal is you can watch it and you can build that trust that the robot’s doing what you want it to do.
27:00 So we’ll see. I have some other opinions about how that might shake out because it’s, you know, it’s not just that behavior model. We have to talk about how to condition that agent level as well. But I also don’t want to leave you with the impression that VR teleyoperated data collection is the only way to do behavior cloning. This is another really fascinating body of work where you have a pyramid.
27:22 At the top of the pyramid, the data is the most expensive to collect but is the highest quality data and will yield the best policy performance. That’s VR teley operation. A big part of that is because there’s no embodiment gap. The data is coming off of the same piece of hardware that the policy is running on. But we have to get clever about these things. You can’t expect to fully scale up data collection at the scale we need it at an internet scale with a bunch of physical humanoids.
27:48 We’d love to sell that many, but it just doesn’t seem like the right way to solve the problem. So, you move down the stack and you say, well, okay, let’s create an embodiment gap and let’s just go around and grip everything with crab claws and we’ll just figure out how to map crab claws to essentially any kind of gripper. And now it’s mobile, now it’s portable. Now I don’t need a whole robot to do it.
28:06 And then you can also say, what if I’m the robot? What if I just put a video camera on my head? And with VLMs today, it will not only kind of understand the environment, the semantics of that environment. It will also semantically understand that this is a human arm and human hands doing work. It will start to create a correspondence between those things. And then you have where everybody really wants to get, which is you just scrape YouTube and every every video frame of every human body doing any type of task that’s ever been put on the internet is fed into a model.
28:42 It’s not that easy. But but you can imagine that like it’s a valuable enough problem to solve that I wouldn’t say it’s inevitable, but it’s it’s probably possible, right? So, it it kind of strikes you like, oh, all of these things that we’re seeing today with language models, like this will be possible with robots. I will be able to talk to a robot. I will be able to have it do whatever I need it to do.
It’s exciting. It’s intimidating, but there’s this reasoning layer at the top, too. So, the agent at the top is a VLM running a particular model. We’ve been working with Gemini, for example. So we’ve used a Gemini embodied reasoning model a lot on a lot of these tests. And then you have the harness which is like AI doesn’t solve everything. So it still needs to talk to other types of software.
29:34 So it’s doing tool calls. So when you’re looking at that you realize oh I don’t need a robot for that piece of the stack at all. Like again I will be the robot. I will tell the agent that I’m the robot. I will give myself a task as a robot and what I’m testing is can this thing take my problem break it down into discrete pieces and ultimately and this is the most important part know that I did the right thing and move on to the next step.
30:04 So you know this is my very loving and supportive but confused wife and I’m like I need a picture for this presentation and I’m like leaning over her shoulders like stack the blocks what are we doing? So you know there’s this future of apps running on embodiment which is essentially I have a bunch of I have a bunch of tasks I have a bunch of definitions of data I need to collect for this large behavior model corpus and I need a system that can assign those out and sometimes it’s going to assign it to a person sometimes a robot sometimes a sim sometimes it’s going to be to flex the behavior level with the behavior clone model sometimes it’s going to be to flex the agent model and then building systems for kind of improving on that.
30:46 And it brings me back to the twin question to kind of like close that out is I used to I used to kind of think of the twin in like three pieces. I was like there’s like reflect, you know, there’s reflect the state of the world, there’s like predict the state of the world largely through simulation tools and then there’s like ultimately controlling the world and kind of creating that that loop.
31:03 But I feel like there’s this kind of funny fourth piece now with VLMs in particular which is the world is an arena to train software and I’ve kind of gotten really fixated on this and I like 3D’s out for me. I like I don’t want to talk about you know classification and segmentation of point clouds. I don’t want to necessarily even do 3D representations or rebuild CAD or reality capture systems.
31:30 I just want to see how far we can kind of push the observance of the world. And well, guess what? We have all these spots out in the world, right? So, we built this thing called site view on top of site map. And you know, this isn’t revolutionary technology. We’re just like, well, these robots are out here anyway. Why don’t we just have them all record video?
31:47 And as we do that, we can have conversations with the factory itself. So, I’m talking about flexing the agent, having a conversation with the agent, kind of figuring out what it understands about the world, what it doesn’t understand about the world. But we can also just be recording more and more data about the world and having a conversation with it. So I used to think that you literally had to measure the world, build up 3D models.
32:10 Now I think you just need a bunch a bunch of video and to keep testing the agent with conversations and kind of building up its understanding. And at the end of the day again this is an economics problem. We are trying to find a way to use the bleeding edge of robotics to kind of fill all these automation gaps in factory through a generally capable humanoid robot.
32:33 Presumably it’s going to be controlled by the factory agent or the factory meases. It’s going to work in tandem with a bunch of traditional types of mobile automation. It won’t just be a bunch of Atlases. Atlases will work with with legacy robots or other Boston Dynamics robots like Spot. And you know this is a vision of the future that we share with Hyundai that we call the softwaredefined factory which again is just like a kind of sexier name for digital twin.
32:56 We just have to kind of change up the lingo every few years. And you can imagine again if we can do all of this right now with software and we have all these methods and it’s a valuable enough problem to solve it seems possible. It seems like we will get there. Part of what I love about working at Boston Dynamics is you get the opportunity to participate in changing people’s perception of what is possible with robots.
33:21 And I think that all of these methods can ultimately be applied to all sorts of other things. Robots in the home, robots in the public space as we work toward a future of robots doing anything. Thank you.
More from CDFAM CD/DC 26

Agentic Engineering: Generative AI in structural applications
Sergey Pigach · CORE studio | Thornton Tomasetti

When Failure Is Not an Option: Bringing Certifiable AI to Engineering Design
Rhushik Matroja · Cognitive Design Systems

From Requirements to Manufacturable Systems: Agentic AI on a Live Engineering Knowledge Graph
Chris Helmerich · Celedon Solutions

The Digital Thread In The Real World: Multiple Partners, Multiple Tools, One Truth
Austin Herrema · Istari Digital







