Podcasts
Paul, Weiss Waking Up With AI
Another Step Forward: Mathematics Breakthroughs and Self-Improving AI
In this episode, Katherine Forrest and Scott Caravello break down how an AI system took on one of the hardest unsolved problems in mathematics. They explain what recursive self-improvement is, how it works, and why it sits at the center of the conversation about superintelligence.
Episode Speakers
Episode Transcript
Katherine Forrest: Hello everyone, and welcome to today's episode of Paul, Weiss Waking Up with AI. I'm Katherine Forrest.
Scott Caravello: And I'm Scott Caravello.
Katherine Forrest: And you know, Scott, the audience doesn't know, but before we got on, we had all kinds of like our usual technical sort of back and forth as for reasons that are known only unto the gods of technical stuff. Things were complicated. But we also have a couple of wild cards on my end that I just wanted to let you know about. We have a dog. I, my dog is wandering around. He's always wandering around, but right now really on the lookout because there's some construction going on in the back of the field at my house and somebody has seen fit to bring a site dog like a dog for the construction site is what I mean. And this cute, very cute dog is a male dog and he is wandering around and sniffing at my dog. And my dog is not into this. And so you may hear gurgling and a dog that sounds like it wants to really just devour somebody. And that may happen while we're on. And the other thing is that they've decided to like cut stone right now. Why? I do not know of all of the times they could be cutting stone. It has to be right now during this half hour increment.
Scott Caravello: I thought the whole purpose of going up to Woodstock was to, you know, get the peace and quiet and get away from all this chaos in the city.
Katherine Forrest: Scott. Scott. Shh! Shh! Shh!
Scott Caravello: Oh, I'm sorry.
Katherine Forrest: It's all peace. It's all peace.
Scott Caravello: OK. OK, all right.
Katherine Forrest: Don't rain on my parade with all my noise. Amy might hear you. And then she wouldn't think it's as peaceful as it is, you know?
Scott Caravello: Right.
Katherine Forrest: In any event, OK, we have a big week here. We've had an incredible week. And we're going to talk about two of the pieces and they really go together. So let me just sort of introduce the two pieces and then we'll start in on the first. But we had first the Navier-Stokes solutionl that was created by an unreleased model from OpenAI. This mathematics problem, this theoretical math problem that we're going to talk about. And for many people, and there's a big debate about this, this is what I call the beginning of the yellow brick road. It's one foot or one hoof onto the yellow brick road of superintelligence. It's only one domain, but it's truly an extraordinary moment. And so that brings us to our second topic, which we're going to talk about, which is right now being discussed in the press around the same Navier-Stokes mathematics problem that's been solved by the AI model. People are talking about, well, what is superintelligence and what is keeping us from superintelligence if it's really going to happen? And there's this phrase which we've mentioned on prior episodes called recursive self-improvement. And that's the idea, idea of AI having the capability to make improvements in itself, which then enable further improvements. And so we'll talk about both those things today, both Navier-Stokes, and about recursive self-improvement, which is sort of the follow-on to that, which is OK if we're entering the era of superintelligence, like at the beginning of the beginning here, but not at the beginning of the beginning of the beginning. So we used to be at the beginning of the beginning of the beginning, but now we're only at the beginning of the beginning. You know where I am, right?
Scott Caravello: I know exactly where you are.
Katherine Forrest: So we'll then we'll talk about recursive self-improvement because that gives folks a little bit of a background into some of what people say is perhaps the next step.
Scott Caravello: Sounds great. So we'll start with the math. I mean, it really has been quite a week. But so this mathematical problem, the Navier-Stokes problem, it's one of seven called the Millennium Prize problems. And so what does that mean? Because I think a lot of people probably aren't familiar with that. I was not familiar with that prize until not too long ago. But back in 2000, a mathematics institute drew up a list of seven problems that are considered the most important open questions in all of math. They were questions that the field had been trying to solve for decades, if not longer. And it put a $1 million bounty on each one.
Katherine Forrest: Right. And that money, which is a lot of money, is really the least interesting part of it because these problems are incredibly complex and they sit underneath entire branches of science and engineering and solving. One does win the prize money, but the payoff, the real payoff, is that the solution and the work that is done on the way to the solution can actually unlock all kinds of answers to unsolved questions generally, and allow for different kinds of devices to be created and problems to be solved that will then in turn solve other problems. And so in the 26 years since the list of these Millennium prizes was published, only a single one of those 7 problems has actually been solved. And so when anyone claims to have solved one of the other 6, it's a really big deal. And so here we've got a really big deal because we now have an AI model that may have solved one of those problems, and that's the Navier-Stokes problem.
Scott Caravello: And, you know, you made the point that it was an unreleased model. And so maybe there's still an opportunity to name it with, like, some sort of Good Will Hunting reference for, you know, solving this like crazy.
Katherine Forrest: What? Does that mean? What does that Good Will Hunting mean?
Scott Caravello: Have you not seen Good Will Hunting?
Katherine Forrest: I haven't seen it, so just sort of clue me. And I'm actually remarkably devoid of cultural references.
Scott Caravello: OK, so like all the action in Good Will Hunting kind of starts when he is a janitor at Harvard and he solves an equation on the blackboard in the hallway that no one had been able to solve. And that kind of like kicks off his whole journey throughout the movie.
Katherine Forrest: Right. OK, OK. And so how does that fit in here?
Scott Caravello: Oh, because the model solved like the unsolvable math problem.
Katherine Forrest: Right. Oh, OK, so they're going to call it Good Will Hunting.
Scott Caravello: Yeah, that was the joke, which is so much less funny now that I've had to explain that entire— wow, that really took the wind out of my sails. But that's OK, moving on, moving on. Right. This equation, it describes how fluids move like water, air, and blood. Navier-Stokes is actually really essential to how artificial hearts function and how blood passing through them in the friction that is happening through the sort of fake arteries. So it's basically Newton's second law, force equals mass times acceleration, but written for a liquid. And nobody has ever actually proved that that equation always behaves under all physical circumstances. So the idea is that let the math run, and at some point it might tell you that the water is moving infinitely fast, which can't happen because water doesn't do that.
Katherine Forrest: So, you know, it's really fascinating what has been said about how some of this was done, which was the deployment of 10,000 agents that were put into groups and that were then allowed to work on this problem and they solved it in 88 hours. So that's an extraordinary feat, but it also shows you the power of the agent. You know, we've been talking a lot about agentic AI and the power of compute, and you put the two together and you've got sort of an extraordinary combination.
Scott Caravello: Exactly. It's not just that it's a powerful system, but it's how long they were able to throw these agents at the problem. And so over the course of those 88 hours, the agents exchanged millions of messages and used 130 billion output tokens. So really, all in all, we're talking about millions of dollars in compute used to solve this problem.
Katherine Forrest: You know, it's probably worth, Scott, just sort of spending a second on what an output token is because people are using that phrase output token all the time. So can you just give a little sort of like sentence or two on what an output token is?
Scott Caravello: Yeah. So we've talked about it a bit when we are talking about the cost efficiency of different models, right? Because it's like by a cost per token. But basically a token is like a unit of what the AI system is actually producing or what it's ingesting. So here we're talking about output tokens like a word could be a few different tokens.
Katherine Forrest: And so you know, what we're talking about is, you know, you've got these— we talked, we talked about like transformer architecture and you sort of you chunk up the data into tokens or you tokenize it. Well, the same thing is occurring on the output side. And this is a way of quantifying the output into tokens. And now metrics are sometimes applied to that in terms of how many output tokens are part of the output or the answer. So all of this is a good bridge because this problem was solved with this, you know, 130 billion output tokens and 10,000 agents in 88 hours. It's a good bridge to our second topic, which is recursive self-improvement, because people have been saying, well, gosh, does the solution to the Navier-Stokes, this particular Navier-Stokes problem, actually mean that we are on this yellow brick road? I happen to personally think that we've got one foot on the yellow brick road. I don't know how many off-ramps there may be. I don't know how long this yellow brick road is. I don't know where it leads to. I don't know if it leads to Oz or someplace else. But the bridge to the second topic is that people say, OK, we're really not going to get to superintelligence until we have this thing called recursive self-improvement. And so what does that mean? And let's talk about that, and I'll start by just saying what the word recursive means. So, recursive is really when something is a version of itself, where you take one thing and it uses itself to make another version of itself, but it might be a different version or a smaller version sort of looks back at itself and then recursively or repetitively improves itself. And so for instance, a recipe, when you make a smaller recipe out of a recipe, the second recipe is recursive to the first. That's just one example, but it's not always about size, it's about something being self-referential in other words. So recursive self-improvement is self-referential improvement of an AI model making another AI model.
Scott Caravello: And so the plain version, right, is that it's AI that's helping build the next AI. But really just to your point, Katherine, it's more than just like AI assisting with tasks related to the AI model's development. And so what we're often talking about for recursive self-improvement is really basically kind of automating this process of creating the next iteration of AI, where AI is designing experiments, it's designing evaluations, it's actually writing the code and it is speeding up the process at a pace that humans could not hope to match. But the whole— sorry.
Katherine Forrest: No, go ahead. I was just so excited, Scott. I just was so excited that I was going to interrupt you, but go ahead.
Scott Caravello: I was just going to say because I think it's just really interesting background is that it's based in an old idea that a British mathematician had put out there in the 1960s and about a system using the intelligence it has to improve what makes it intelligent, which would lead to a, quote, intelligence explosion. And so that's what we're talking about, right? Because that pace just kind of snowballs once you have more advanced AI creating even more advanced AI. And the pace is just exponential.
Katherine Forrest: Right. And that's really the point that I wanted to focus on. So the concept of recursive self-improvement, when you're talking about it in the AI context, is that you have this moment when AI is able to improve itself and is able to do that at such a speed that it no longer needs humans to do the improvements. And it can do improvements as a result and at a pace that are beyond humans to control, to stop, or even potentially to understand. We don't know because we haven't gotten there yet. And so one of the measurements that you see in system cards all the time now, and this is relatively recent, is if you look at model cards or system cards for different models, you'll see that there's a section usually on how far along the road towards recursive self-improvement the model is. Now there's one thing that I wanted to go back to and we covered it in an earlier episode, and it's a piece of jargon which can be relevant here, which is called a harness. And I think we really should do a whole episode on a harness.
Scott Caravello: Completely agree.
Katherine Forrest: And so, you know, if the model is an engine, then the harness is the rest of the car, it's the steering, the gearbox, the brakes. It's the software around the model that decides how it plans, what tools it reaches, you know, how it checks its work. Almost all, virtually all of the self-improvement of AI models that's happening today isn't recursive self-improvement, it's really human directed. And so even while you know that you might have AI writing some of the code, it's human directed.
Scott Caravello: And so we've said that we're not there yet, but how far along are we in the actual process? What's actually happening at these companies building advanced AI?
Katherine Forrest: Well, you know, we've got now a situation where there have been models where it's reported at least that they've written 80% of their own code. But again, it's human directed. They're actually prompted to do so and prompted to validate sometimes, sometimes agentically, they validate as part of their overall process. You can do a single prompt that can actually prompt both the writing of the code, the validation, the patching of the code, and then the rerunning of the code. But we're not yet at a point when, you know, we're able to say that the model is just prompting itself.
Scott Caravello: So branching off from that and what do people see as the risk? Because recursive self-improvement has been all over the news lately because it's tied to what people see as, you know, some of the pitfalls when it comes to ever advancing AI technology. And that idea starts really simply. If an AI can do all the AI research, writing the code, running the experiments, designing the next model, and like we were saying, accelerating at an exponential pace of improvement, then basically we just get to a point where the machine stops waiting for us and we don't really know exactly what's going on.
Katherine Forrest: Yes. The concern is that there might be a dangerous loop that could lead to a loss of control, and we just don't know. We don't know if that's going to be what happens or whether or not recursive self-improvement could end up unlocking a huge amount of advance in AI tools, AI models generally. But that is the concern. And so people take the recursive self-improvement sort of phrase and they say, OK, if we hit recursive self-improvement, boy, the world could be a really different place. But nobody knows what that place is going to look like. And so, like I said, we may have a foot on the yellow brick road. We may not have a foot on the— well, yeah, I think we have a foot on the yellow brick road, but I don't know. Do you think we have a foot on the yellow brick road?
Scott Caravello: I think it's becoming more and more likely that we have a foot on the yellow brick road.
Katherine Forrest: Yeah, it's really interesting because for a long time, you know, and I'm partial to this because my book, right, my book, my book was really timely, right? For those audience members who are listening to this for the first time, it's called Of Another Mind and it talks about superintelligence. But anyway, putting that aside, my book was all about sort of, is there going to be superintelligence? There will be superintelligence. It's coming. When it comes, we don't know what it's going to look like here at different scenarios. And now we're actually at a point where not only is it coming, but we may have that foot actually on the ground on that yellow brick road. So, you know, I think that's all we've got time for today. Scott, you want to have any closing remarks on all this?
Scott Caravello: You know, I guess I would just sort of close out on one caveat about the concept generally, which is just about the bottlenecks that it might hit in the real world for AI to continue advancing. We've been needing ever larger amounts of compute. And so that's limited by what's actually out there and data center construction and what these labs can actually make use of. And so that's often cited as a potential limitation for the time being on AI actually reaching this point where it is recursively self-improving.
Katherine Forrest: All right, so TBD, right? We're going to be watching this. We're going to be watching and seeing when we get to the point where somebody claims we've got recursive self-improvement. So that's the end of our episode today. Folks, there is so much happening right now in the AI world. If you've got a topic that you want us to cover, please do write us and let us know. Otherwise, we're going to just pick our favorite topics as we do each week. If there's a good topic that somebody comes up with, we'll send them one of our Waking Up with AI mugs. Because I'm long in the AI mugs, I've got a lot of extra AI mugs Waking Up with AI mugs, so we'll do that.
Scott Caravello: I think that's great. You know, we're moving off so soon and I have two boxes myself that I would love to start, you know, distributing before they have to come with me. So this is a great contest.
Katherine Forrest: OK, so we need a lot of ideas. All right, now we've got a lot of ideas ourselves, tons of ideas, because there's so much happening right now. But if you send in one, we'll try to use it. All right, folks, I'm Katherine Forrest.
Scott Caravello: And I'm Scott Caravello. Don't forget to like and subscribe.