CDFAM NYC 2025 · New York · 29 October 2025
Superintelligence for scientific discovery in the material world
Abstract
AI is rapidly transitioning from a passive analytical assistant to an active, self-improving partner in scientific discovery.
In the material world, this shift means developing systems that not only recognize patterns but also reason, hypothesize, and autonomously explore new ideas for design, discovery and manufacturing.
This talk presents emerging approaches toward ‘superintelligent’ discovery engines -integrating reinforcement learning, graph-based reasoning, and physics-informed neural architectures with generative models capable of cross-domain synthesis.
We explore multi-agent systems inspired by collective intelligence in nature, enabling continuous self-evolution as they solve problems.
Case studies from materials science, engineering and biology illustrate how these systems can uncover hidden structure-property relationships, design novel materials, and accelerate innovations in medicine, food, and agriculture.
These advances chart a path toward AI that actively expands the boundaries of human knowledge in engineering.
Transcript
From the speaker’s corrected captions. Each timestamp opens the video at that moment.
Read the full transcript · 5,956 words
0:00 So, good morning everyone and thanks Duann for organizing this. Very nice to see this really really nice crowd here. So well it’s kind of hard to be the first presenter. I don’t see all the other ones but I’m going to do my best to kind of set the stage for the next few days and I will talk about you know, really some foundational questions I’ll be raising and some, you know, remarks, some thoughts on the future of discovery.
0:19 As you can tell from the title, we’re trying to figure out how do we build, you know, really intelligent AI, you know, no pun intended, that actually can teach us something new. And, you know, as you know, I mean, if you look at the sort of the history of how we have that’s okay, I’ll use a manual clicker here. The the history of how the world sort of came about.
0:40 I mean, you can think about it as sort of creating the material world itself. You know, we had you know, sort of evolved as a species to to make use of the world. I guess building tools and machinery and things like this. And I think what we’re really right now witnessing is that we’re able to create machines, you know, systems, synthetic constructs that actually, you know, can can imagine themselves and they can also design new things themselves.
1:10 Okay? And they can do this already today. You know, you think about AI writing code and executing the code, you know, programming a robot, experiencing the world and and really sort of, you know, creating through this process innovation and discovery. And that’s the main thesis I think what I’m trying to convey is, you know, there’s obviously many aspects to this and we won’t talk about all of them, but I’d like to take a very positive look on this is if we can figure out how to automate discovery and make it make it abundant, we can solve any problem we want.
1:37 Anyone can, in fact, and that’s a really empowering kind of future I think we can see there. As you know individuals we’ll be able to do things just we couldn’t do before and it gives us sort of this superpower and you know of course that’s what we want to be right and so the the first thing I want to talk about a little bit is you know where are we today and and what are we seeing today and I think if you and you probably hear about this in this conference maybe more from an engineering perspective but you know broadly in AI I mean you’re seeing a lot of headlines right like this you know this is more about 10 years ago you AI beating the best goal player in the world.
2:13 On the right hand side, this is a news from this year, AI winning gold medals in math. Okay. And then you see everyday headlines like this, you know, you know, AI is beating the benchmarks or saturating the benchmarks and stuff. And and it sounds really cool, but as we probably all know in this room, and maybe maybe if you don’t, then maybe you can think about it.
2:33 You know, these are all, you know, really impressive things, but they’re really not discovery problems. They’re not open-ended. You know many of them are close problems you know like math you can you know essentially you can you can test it right away of course scientific discovery is very different right or innovation or engineering it’s not something that you know exists or have never been built you know in engineering something you imagine could be built and then you have to kind of figure out how to do that but AI is not very good at that so it’s really really good retrieving things okay and and and that’s really what we are so the the things we do in the lab spend a lot time probably most of our time trying to figure out how do we build you know AI or models systems that can actually teach us something new what are the principles there and you know a few things I always say is you know people always talk about data and I always make the counterpoint and say well data isn’t really everything in fact you know we’re never going to have enough data data alone cannot bring us out of this problem you know that we can’t really create discovery I think we we can you know produce more and more data like in biology or manufacturing or maybe other things, sensors in cities and things like this.
3:41 But it doesn’t really help us sort of really solve the problem I think we really got to do and that is you know the two components I would say really of of human intelligence and that doesn’t mean we’re going to build synthetic intelligence like us but but at least what makes us different why can we innovate right why can we you know be scientists and discover new things and one there are two components right one is we’re able to reduce complexity principles like prototypically would be you know finding an equation Right?
4:13 And then and then using the equation to do compositional reasoning and so say oh I know I know a principle and I can solve a new boundary condition that’s a very simple one or I can come up with you know in art you know creation of new ideas you know I can combine principles in really novel ways and I can do it in a reasoned way not just randomly actually you know reason so so those are two things I think we’re missing in today’s AI.
4:33 I made a little actually picture this morning hopefully is more clear you know so for today’s AI really we’re just learning from the data and we’re approximating the data right you know that’s the sort of the easy path right and and you know it’s not enough you know we got to create these this tension here between reductionist and compositional principles it’s much harder for something to learn that a machine to learn this but you know ultimately something like this is necessary to really go beyond the training distribution so you hear me say this a lot.
5:04 So I’ll show you many examples as well, but this is sort of how the world looks like to me. And you know, in in in a way, you know, the problem spaces we’re interested in, like Ran said, we’re looking at materials. They are really complex systems. They’re hierarchical, but what it means is they’re not just hierarchical in a sense that they, you know, have atoms and molecules and other building blocks.
5:22 They actually have relationships, interactions, across scales that are non not obvious, non-trivial, right? You might have an atom, a single atom that’s really important for the function of something and you know other atoms don’t matter, right? So, so there’s not just simply harical structuring like we usually think about it, but actually it’s very complex. And so for the early days, you know, at the 1950s computational modeling, we we’re trying to basically build models atom by atom, you know, trying to figure out how do we, you know, maybe simulate our way out of this.
5:52 And, you know, I’ve gotten involved in this in 2000, right? You know when I did my PhD and I and I you know we thought you know this is a potential very powerful right but of course you can never again scale your way out of this because you’re just brute force computing maybe motions of molecules and atoms it’s a powerful tool but really doesn’t get you really to the the meta learning and that’s what I got interested in in about 2010 I became interested in this because I realized that oh you know the way I think about a problem when I give a talk about let’s say spider silk I’m not telling you about the motion of every single atom right and then I’m just going to spend an reading all the the dynamical trajectories of these.
6:28 No, I’m actually going to tell you about principles. I’m going to tell you about scaling laws and motifs maybe you know certain you know relationship that are really critical and I and I and I give you kind of a an informed perspective and so we’re trying to figure out how can we model this actually right this discovery process you know and and that’s sort of kept us really busy so we did some early work in 2010 using mathematical concepts I called category theory and I’ll show more on this later to kind of get this kind of stuff this is an actual spider web being created by a spider And you know there’s no hope.
7:02 There’s no law here. So you know can we not just reproduce this in a model but actually understand the principles. How did the spider do this? What decisions the spider make? And why do these structures come about? So this is an actual scan of a spider built by the spider and it’s you know fascinating but a complexity right it’s very very high. So, you know, I think when I I always like to think about sort of the again to make this point again what’s different from brute force or machine learning or even just modeling by writing code and letting computers execute the code is that you know we’re very good in abstractions right we we we can do logic right and there’s many examples you know we also figured out I think in you know maybe a few hundred years ago at least maybe before but few years ago formerly that you know in science you really need to look at differences you need to understand what are the things that are shared shared principles what are things that are different right between different systems and so this really is how we do science and that’s how we discover the world and you know a lot of the early I would say kind of work you know in in in neural networks machine learning totally ignore that right that principle because they basically approximate the world as a closed system and they you imagine the world can be modeled.
8:19 But of course discovery you know requires an outside perspective right it’s not enough to just have an internal viewpoint and so this is something that we we are you know thinking very hard about how do we do this and one way you know to formally mathematically ground the work is is in sort of the math of math which is category theory so as I mentioned earlier so this is how we started doing this in 2010 and there’s obviously not enough time to really go into everything I’m probably behind already in what I wanted to say but the point is that you want to kind of have a formalism where you understand you know isomorphisms that is things that are shared between very different systems and what are the mappings between them right between let’s say music and structure materials and language and other things and back in those days we did this manually meaning sort of by hand you know writing analytical equations and you know functors and relationships and things like this and it was very successful because I could basically condense the things I would give a talk about about spy silk in into a graph right and I can actually sort of capture this very sensual relationship in a effective way and we’ll come back to this later you know you have to compress the information right in a very high high density of of expression and you know coming back to this I think if you think about sort of differentiation sort of the finding differences and finding similarities is really I think most beautifully expressed in differential equations that was the early the first time I think humans really figured that out and I think we didn’t figure it out actually at that time that this is a principle we just find found a way of writing the behavior of systems, dynamical systems, the world with these differential operators and you know we have you know kind of now I think opportunity to go back to this actually as we model intelligence itself.
9:58 So that’s the the thesis behind what we’re doing and of course when you do this you have a couple of things you need to kind of unlearn I think one is you you can’t really train a model right on on data and then expect it to discover things outside of the training distribution. It’s maybe obvious but not obvious to some actually. I think there’s a lot of effort in in you know put into trying to do that and I think it’s it’s not it will it cannot work.
10:23 So you have to have machines systems models whatever you call it that that can can learn obviously on the fly you know they can adapt and I think that’s the consensus I think we’re getting as a community now. That means you’re you do some training and giving models some basic intelligence maybe but then you also let them evolve and they actually really learn about the world doing the solution of the problem right kind of like humans except if we use you know you know reinforcement learning techniques for example we don’t have to you know bias these machines to do what we think should be done they’re going to discover it on their own right that’s that’s the idea behind it so this is I think the the framework for a lot of the stuff we’ve been doing and of because nothing’s ever new.
11:06 That’s part of category theory, right? In that you know the in fact the early ideas you know how people thought oh let’s build AI in the 1950s and60s was exactly this idea right and so as you some of you who how many of you have been around back then I haven’t been around anyone from this time probably not okay but you know back then it was a great idea right and it didn’t work it basically became humans writing programs if and then statements literally right and and then they and they would break as soon as something was slightly different so it took another 50 years to actually, you know, create these sort of symbolic methods and connect and marry them with neuronet networks which provides a connective glue.
11:49 It’s like a a glue connective tissue, right? So, so it’s the the relationship becomes you know, becomes differentiable if you wish, you know, pun intended. And and you can actually make these connections even in edge cases because neural networks have an ability to extrapolate maybe some and they can interpolate, right? And so they can make connections. And so this allowed us I think to build these systems now like this is a model we’ve built for protein design and the many others in this field and and sort of the idea of having sort of self-organizing you know problem solution strategies you know that learn on the fly right they learn from things they they try the experience and it doesn’t mean the experience has to be in the physical world it could be in silicon entirely it’s fine but they have to sort of experience the world and update their beliefs biology does it the same way right so biology ology obviously you know creates these evolving systems and to give you kind of how many of you work in biioaterials or anything very few okay but if you yes if you’re bio you know you know that biology doesn’t engineer a new body every single time a human is born right I mean it’s it’s it’s basically reuse principles even our organs are made from the same stuff chemically you know they just evolve to you know create new functions by putting the building blocks together slightly differently and I think an engineering we’re kind of want to do this but it’s hard to abstract that out and figure this out.
13:08 Generally though I mean I would say as a species we we are not doing this really well right we’re highly inefficient in the sense that we’re creating thousands of different materials for thousands of different functions as opposed to sort of using what nature does and intelligence is actually the same way. So the same story you know we’re building gigantic you know some people call it foundation models or whatnot but you know sort of trying to make this monolithic thing that is sort of very good in solving the problem but actually that’s not how you know we can create most intelligence materials or most intelligent intelligence itself probably and of course there examples in biology like bones and growth and things like this and actually this is a slide from my student Lee who’s right here so you know something that Lee’s actually interested in is figuring out how to extract some of those biological principles and how design works in nature or in biology and then creating sort of a whole pipeline of of extracting these principles and then being able to fabricate them.
14:05 So I thought this is something interesting. So Le here you can chat with her during the next few days how this can be done. It can be end to end potentially right. So you can have you have a machine that would you know study nature and explore principles and ultimately come back and and build something new that is compositionally created. So it extracts the principles like we show earlier and then it uses the principles to compose something totally new with a certain objective right and it does it through thinking and reasoning and in a more abstract way we call it compositional reasoning.
14:31 So this is sort of the the terminology technical term for this where you have different components and they come together and form you know you know an answer. So there are many ways we can sort of approach this. One is thinking okay what if I take you know hundreds of you know intelligent agents and they come together and form a swarm right I’ll show you this later yeah that can work you know you can have maybe reasonably intelligent systems that have some understanding of the world and then they they act in the world and they express ideas you know they imagine things they test them and they learn they can update their beliefs about how the world works that’s that’s important so they don’t just come in and and know everything which is a current paradigm but they become part of the solution, right?
15:14 And they evolve like insects, bees, and things. They, you know, they they’re not born to build bridges. They, in fact, none of these elements actually, right, know that they’re building a bridge. They just figure out how to do it through trial and error and adapting their own behaviors. And so you know as a statistical mechanician if you remember sometimes I call myself that. Because I come from stat mac in molecular simulation the world really becomes oh I’m not just optimizing the particle I’m optimizing the interaction right so so there’s the particle property like what atoms do I put in but you know if I want a certain structure I really have to engineer the interactions and the emergent interaction so it becomes like a problem like how do I figure out how the same atoms interact slightly differently in creating a very different outcome right so this is sort of the design problem so when you think about training a model in quotation mark like you have to train them to achieve the set outcome.
16:11 And again, we we want to do two components here. One is we want to kind of get the models to learn shared abstractions. So again, the idea of science being understanding differences and similarities, right? Which I think is the core of scientific inquiry and and discovery. And you got to then kind of you know use this this you know this this structure to go outside the training distribution, right?
16:36 That’s that’s the core. And if you want to go one level deeper, how we do it? Well, we do it through graph representations like I showed you with the categories. So we we kind of think about not just memorizing the world as it is or or even using graphs that you might know, but discovering the graphs on the fly. So it’s a key different, right? So if you think about a molecule and or a social network and you were to say, oh, I’m going to put a graph on the friendships, right?
17:00 Well, that’s you telling the model that’s what I think the relationships are. But the you know the latent relationships might be very different right there might be certain features and traits that are actually more important than others. And so what we really got to do we got to build something that can discover the graph and then we compute on the graph right and so we do both of these things and if we do this we can you know build sort of at least individual components that are you know pretty good in approximating the world like this is work by my student Jamie who who’s been really fascinated with cell automa.
17:28 How many of you worked in cell? I thought many some of you do. Yeah. So those are kind of really interesting systems. They’re like the the simplest possible algorithm I guess to describe evolution of a dynamical system if you wish in that or design maybe and the experiment was well let’s use some of those models and neural network type approach and let’s see if it can approximate can learn these rules and it can through an approximation right so you saw from the title I said it’s an approximation in fact that’s all it is right so we can learn the we can never really learn the rule fully right we can approximate the rule and we always find cases that are never been in the training data.
18:11 And so now you can show and this is the point of that paper actually was to basically show that yeah you can approximate the world even for very very simple systems but you can never really understand it and never fully make predictions outside and I think it’s an important argument to make actually in the community. Then we thought well what about using some of the symbolic methods like mentioned earlier right?
18:33 So what if we were to use this kind of framework but instead of just teaching the model to approximate the world directly can we you know use a model like this to tell us the rules that govern the world right that would be kind of extracting a symbolic representation of the dynamics and that works okay I mean we can show in in this so we can do forward and inverse problems so you give it a evolution of a system right a new frame and the model will tell you this is what I think is the best dynamical you know rule rule to predict that next state in the system.
19:05 And the more rules I use for training, the better performance becomes. And at some point actually, and you can see this on the very right here, the model is able to, you know, infer rules that were that have the same or similar behavior, right, as what they see, but they’re different from the training data. So there’s some sort of interpolation going on in in this system. And so now you’re beginning to have something that you can put into into a a more sophisticated learning strategy based on reinforcement learning.
19:39 Now you have an analytical expression of how the world evolves and you can probe it and you can test it. You can see is this actually true. And this is what the differentiation means. So I find a rule and I can now form hypothesis essentially based on this mechanism. It’s a actual hypothesis not just a prediction you know from a black box but it’s the actual analytical expression.
19:54 So that’s sort of the spirit of this work. So what do we do next? Well I mean basically we want to go beyond this right? So we want to kind of figure out how do we go outside the the training data and you know one of the things we got to do is to use better training strategies. So instead of teaching the model through examples we want to teach the model based on outcomes right and so we say you know I’m going to reward you if you figure out a good you know strategy to get to the answer and I’m going to penalize you if you don’t.
20:27 And so that’s very different from what we call supervised fine tuning which is most trades. You need a lot of data for this for reinforcement learning actually you don’t need a lot of data right because you’re really teaching the model strategies and if a model can do this well through abstractions it can do a pretty good job. So that’s sort of the the approach we’re doing a lot in these agendic systems.
20:45 So there’s different ways to get there. One is in fact sort of a multi- aent system. You know many AIs talking to each other and negotiating each developing its own strategy each becoming a policy right to to solve a problem and and this you know is is fundamentally different from sort of traditional learning I mean this is one of the first papers that we did this in 23 and I basically you know I I gave this model a task and I said you know can you give me this energy landscape of this molecule and if you were to take a traditional I don’t machine learning potential or maybe like traditional neuronet network would have given me an answer even though this the model has never seen this this particular molecule in his training data.
21:31 This model did something different. So I I put together basically a system of agents and it figured out that I I doesn’t know the answer but it knows how to write a code right that runs a DFT simulation to generate new data to then do the regression problem right so this sort of was eye opening for me and and I mean I’ve actually tried this for many years and it could never work right and you know at some point in 23 we had enough sort of baseline intelligent models that could do this and it changed sort of the way we can think about problems Because now you know the machine actually it writes a program a code.
22:04 It it solicits new data. It has agency in the world. It can search the web you know the internet and things like this. And it can test hypotheses and then then maybe what what can we do with this? Well, so this was a toy problem in the beginning. We went on and did many different things with this. We built real multi- aent systems meaning real that actually solve important problems right you know like protein design and then with the amazing thing is you know suddenly first of all the model can make its own data so you don’t just train it and then you’re done and you hope for the best but hey I can actually the model can run physics simulations it can you know search the internet it doesn’t just you know the world through neural networks like a large language model is like seeing the world in a very fuzzy way right so so it just sees the the mundane like average truths, right?
22:54 Which are the most established. And of course, there are edges in the in that world, you know, maybe data that’s an outlier that actually is the key to make a discovery or you you know, you have to be able to you know, have you know, conflicting representations of the world. So neural networks are not very good at this because they’re trained essentially to give you just the most likely outcome.
So and and so this allowed us to change this. We also the other really big thing here was that suddenly we we were not just doing datadriven learning or physics we could combine them right this model had a physics engine it could run a simulation and what I mean by this is the model actually it would write its own code to run the simulation it wasn’t me telling the model here’s how you do it but actually could execute on its own so it write the program execute program collect the data and do things and you can do amazing things with this so you can solve mechanics problems design problems like here.
23:52 The the burden of you know sort of finding the perfect answer is gone because you know the the system can negotiate solutions and error correct right so even if the answer initially is wrong the code doesn’t run you know there’s context the model can understand how to fix it and it can right so so there’s this connective glue actually that makes these things work and it can ultimately solve the problem through many many iterations so so you might say well what’s the catch here one catch is compute.
24:21 So Nvidia will be very happy. I think we have some video folks here. So, so it takes a lot of compute but it means actually that you need less training because I can use this agentic system and I can solve this continuum mechanics problem today and then maybe this is another slide from from from Lee you know we can do an add manufacturing design problem tomorrow right so so it’s very flexible some of the students in my group have done work on applying these systems to binds by design this is work on you know using pollen particles in the specific domains we’re connecting sort of a baseline intelligent model with a very specific domain with experimental partners.
24:57 This my student Rachel did this work and then trying to figure out okay can we can we use agents like this to come up with creative new designs and creative new ways to actually conduct an experiment and then go to the lab okay and then actually you know do this experiment there and of course there are many components and I’m just sort of flashing a couple of key ideas here you know you have to create the model right the 3D architecture you have to design the experimental plan the models have to be able to right?
25:24 They have to look at results and reason over them, right? Connect the dots between different concepts and ultimately you know they they can they can they can execute you know things internally in silicone that they can then be put into a lab and actually validate you know some of the design ideas and suggestions that come up. So this is very exciting. So you know it takes a lot of work though.
25:51 I mean in in this work we spend as much time maybe more in the experimental side of things. This is experiments done by Nam Chucho who’s an expert in in this particular so it’s a very particular field right plant pollen based materials and why that is somebody I know and I like him and we sort of formed the team together. But that’s how you can really test sort of how these creative AI agent systems actually are are working.
26:12 The other thing that we going to do when you do things like this you know you can do human experimentation but of course you can also do lab experimentation autonomy. So this is something also we’ve been interested in is to build you know hardware that can be used you know for autonomous collection of data right so this is also something a lot of people are doing right now and one of the really important things here I think is is open sourcing these things so this printer we we made this work by former student Jesse so he you know created this printer built it and you can see a little bit in the flowchart there’s also paper which you can reference there and it is open source and it’s not just the code but the entire manufacturing.
26:53 I think that’s really important for the community to figure out how so you you can take this and make it improve on it put it back on GitHub right and you know or make your own GitHub repository and then the community can involve it’s very important I think to have this in software and AI machine learning right so PyTorch TensorFlow and the whole machine around it the reason we’re going so quickly in this is because we have a shared language we speak you know we can all talk to each other I was saying yesterday right I can read a code that an AI lab has created and I can adapt this go to my own purpose so I can improve on it.
27:26 But in engineering a lot of times not the case. You know we can’t really talk to each other very easily. So this this is very important. So and hopefully maybe this community can sort of get excited about this. There many other examples of course how we’ve been you know building you know solving problems using multi- aent systems. One is a protein design. That’s something we’ve done a lot of work in you know designing proteins with these systems again including genai reasoning and physics right so these are not just predictions they actually have a physical meaning because they’ve been validated so they have been molecular simulations mechanics to actually give us a signal whether a design actually makes sense or not right and so you can do this to whatever degree of perfection you want we can also do this for different materials we’ve done agent systems for alloy design paper in PNAS you can check it out and we’ve done this even for scientific discovery so the title of the talk really was about scientific discovery framing it but you know and and I mean absolutely you can you can automate the entire scientific process now that we can do and what I mean by this is not just to create a design solution which is not I mean just in quotation mark it’s very hard but to actually come up with principles so can the model teach us something about how the system works Thank you
More from CDFAM NYC 2025

Real-Time Computer-Aided Optimization (CAO): How GPU-Native CFD Changes the Industry
Gregory Roberts · FlexCompute

Design You Can Trust: Explainability and Control in Physics-Driven Generative Design
Marco Pietropaoli · ToffeeX

Shaping Flow: Computational Design Strategies for High-Performance Liquid Heat Exchangers
Ryan O’Hara · Alloy Enterprises

Accelerating Metal-to-Plastic Conversion with AI, Implicit CAD, and Mesh-Free Simulation
Karthik Rajan Venkatesan; Neel Kumar · Eaton; Intact Solutions

Computational Craft: One Footwear Designer’s Quest to Replace Himself
Samuel Whitworth · New Balance

Acoustic-Driven Computational Design: Premium Branded Audio in the Automotive Industry
Austin Mitchell · Harman International

Computational Morphogenesis: Leveraging Proceduralism to Unlock Temporal Design
David Burpee · David Burpee

Engineering Intelligence: Practical Applications of AI in Structural Engineering Practice
Sergey Pigach · CORE studio | Thornton Tomasetti

















