Podcasts
Paul, Weiss Waking Up With AI
Recursive Self-Improvement and the Road to Superintelligence
In this episode, Katherine Forrest discusses recently disclosed incidents in which AI models operated beyond the confines of their testing environments, accessing systems belonging to outside organizations. She also explores recursive self-improvement, the "Pacing the Frontier" statement, and proposed federal legislation aimed at preserving human control over advanced AI systems.
For the sources referenced in this episode, please see the links below:
Hugging Face: Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident
Pacing the Frontier: A statement from 1,350 employees of frontier AI companies
U.S. House of Representatives – Office of Rep. Ted Lieu: AI Kill Switch Act
Episode Speakers
Episode Transcript
Katherine Forrest: Hello. Folks, welcome to today's episode of Paul, Weiss Waking Up with AI. I am Katherine Forrest and I am solo again, right now, just for a day, Scott is off busy doing something and he's in some undisclosed location. And actually I think he's in a very disclosed location. I think he's sitting in his office, but he's actually caught up in something so he can't join us today. So, I told him I would take this topic today that we had, which actually is near and dear to my heart relating to superintelligence and some models, and go ahead and do it solo. And you're getting me, by the way. You're getting me in Maine, which, you know, I've been here before, you guys have been with me, been here before. And I just want to tell you I have to put in a plug for this because this was so amazing. So I just ate a burger from a place called Higgins Beach Market. Now, those of you who know the area of like Cape Elizabeth and Portland and Scarborough, you have probably heard of Higgins Beach Market, but if you haven't, it's in Scarborough. It's really great. They've got great lobster rolls and great, which I do not eat because I'm not like a lobster person, but a lot of people tell me that they're fantastic. But this burger, this burger was really great anyway, so I just wanted to sort of tell you because it makes me ready to go–ready to go. And the other thing that makes me ready to go was I had this amazing experience where in my backyard I had a great blue heron and it was—they're first of all, they're incredibly tall. OK, who knew that herons, I'm pronouncing it right, but herons, herons are incredibly tall. They're like 4 feet tall and they're beautiful and anyway it was sort of standing there looking for its breakfast. And so it was extraordinary. So those two things today mean that today is a good day and we're ready to go, OK.
And so let's get right into it, so I want to mention what has been coming up more and more actually over the last not just this week but couple weeks, but in a cadence this week that has been extraordinary to me, which is superintelligence and many of you who listen to this podcast regularly know that my book on superintelligence came out relatively recently, Of Another Mind, which you can get at Barnes & Noble or at Amazon, but it's about superintelligence and navigating a new social contract with superintelligent AI. And so I co-authored that with Amy Zimmerman and it happens to actually now be fitting right into the conversation that is all over the place about superintelligence in the last, really in the last couple of weeks there have been articles, interviews with folks. There actually been a couple of interviews with people from Anthropic, some current and former security folks, safety folks, and scientists, engineers talking about superintelligence and what it could mean. And I watched one of those a couple of days ago, two days ago, from, which was done by a guy named Jeffrey Ladish, L-A-D-I-S-H, which I recommend people watching. You can get it on YouTube and the debate really has moved on from will superintelligence ever be possible to when, and people are talking about as early as 2027. I'm not sure I'm there yet, although as we were saying, you know, there's kind of a proto-superintelligence already, to 2030 and what's it going to look like. And so these are all things that people are talking about and one of the debates which has come up in the past week with some of the new capabilities of the newest models that we talked about some during our last episode is AI actually improving itself, and that when AI can really improve itself in a robust way, a 360-degree way, manner, that we're going to potentially have a new jump in intelligence and possibly the jump that we are expecting into superintelligence.
And there is something called “recursive self-improvement,” which is a kind of self-improvement that AI could itself engage in, and when it has improved itself, it goes back—that's the recursive part—and it improves itself again. So it's sort of a loop if you will, and that's the recursive part, recursive self-improvement, this feedback loop. And AI would then be able to, as we know it can already write code, but it would write the code, it would test the code, which it can already do. It would develop new capabilities, potentially. It could even potentially develop new architectures and port over the code into new architectures, discover new efficiencies, propose and implement successor systems, all kinds of things.
So this is something that people are really thinking a lot about, and one of the reasons is that we know that the Claude Mythos, the newest Anthropic model, Opus 5, they've been actually engaged in writing a lot of the code that is at their base. So this has become sort of a big conversation now corresponding with this conversation on superintelligence. And not by chance, there was this week a group of 1,300 employees from frontier AI companies, and these are people who are among the highest-level engineers and scientists at the most recognizable names of AI developers. They put out a statement called, quote, "Pacing," P-A-C-I-N-G, "Pacing the Frontier," and the statement, which is very short and you can see it on the Internet, the statement says that the leading companies may be close to actually automating AI research, which is in part the recursive self-improvement, and warn that AI capabilities might soon move beyond our human ability to control them. So the statement again is called "Pacing the Frontier" and it's really worth taking a look at and go into the who signed and look also at the comments because the comments are actually signed comments from some of the signatories themselves and so you can see what X person or Y person is actually saying in particular, and those are also very interesting. But the statement, it suggests that companies, the government, and society might actually need more time to address what are emerging risks, and they ask the US government to support an international effort to develop technical and governance tools that are needed to deliberately pace—hence the name of the statement, "Pacing the Frontier"—to pace the frontier of automated AI development, again crossing into recursive self-improvement. Essentially saying we need time so let's slow this down. So take a look at it. Really interesting.
But there's another side of this, of course, because the US and the Western countries are not the only ones who are involved in sophisticated AI model development. We've got China that now with Kimi K3 is considered to have one of the most sophisticated, and that's an open-weight model, around. We don't know what China might actually have that has not been released outside of China. But at least Kimi K3, which has been released outside of China, is extraordinarily capable. So this effort of pacing the frontier runs into how do you get that kind of international coalition to do that. So there's some things we'll be watching there. But the reason that superintelligence is so much in the news is, you know, about control, about the concern of what's going to happen when it arrives.
And some of that discussion has been spurred on because of the recent sort of news articles that many of you listeners may have seen about some incredible capabilities, including cyber capabilities of some of the most recent highest-capability models, and some of those models actually having an ability to—using the word hack—hack into or get out of secure environments. Sometimes, you know, in most instances, but not all, these have been part of tests that companies have run, and sometimes models have actually exceeded actually achieving what they were asked to achieve by leaving an environment and thinking that, you know, they were supposed to leave the environment. But they were in fact exploiting a vulnerability that people didn't know existed. So, the testing itself is becoming its own thing. I'm recording this right now on July 31st. You'll be hearing it a week from now, and those of you who are listening to these things asynchronously are hearing about it at some other time.
Now let me jump from the Kimi conversation to something else that's been getting a lot of attention involving a model hosted by Hugging Face that's called GLM 5.2. And those of you who follow regularly may know that we've talked about GLM 5.2 relatively recently because it's made by a Chinese company called Zifu AI, or known as Z.AI. That's how they're known globally. And it's a very capable open-weight model. It's about a 750-billion-parameter model that was released in June. And two weeks ago, I think, is when I sort of had mentioned it. But here's where this GLM 5.2 comes in. And this is the part that I find really fascinating in light of this discussion today, because Hugging Face used GLM 5.2 in its response to the intrusion to figure out actually what happened. And it did that by analyzing over 17,000 logged events from the breach. And by that, what I mean are digital records of actions and things that had happened that were actually recorded on the Hugging Face systems. And those were all recorded in something called logs. And logs in sort of cybersecurity lingo is sort of an audit trail, if you will, about what's happening on your network. And if something goes wrong, you can then try to go and figure it out and investigate it.
So, what GLM 5.2 actually did in that forensic effort is really striking, that really does highlight its capabilities because, according to Hugging Face's write-up of events, which, by the way, I suggest everybody read, it's really interesting and publicly available. I keep saying really interesting, and that's because there are so many things that are really interesting. But there was a general automated scan of logs that didn't reveal what had happened. But the GLM 5.2 was able to reverse-engineer the encryption processes that the intruding model had actually used to hide its activity. So by this reverse-engineering process, Hugging Face was able to figure out what happened. So when you couple that with the Kimi K3 capabilities that we were just talking about, you start to see why the testing, the guardrail design, the whole conversation about, you know, how we build safety into these systems becomes even more urgent and more complicated than people think.
Oh, that's my dog barking. But you can sort of all hear that dog because the dog actually believes—this dog strongly believes that—if you can hear the barking in the background, I don't know if you can or not, but if you can't, I will tell you that there's a dog barking in the background. This dog strongly believes, and it's been passed down for generations, I think, through breeds of dogs, that the mailman and the FedEx person, the mail person and the FedEx person, are his arch enemies. I don't know why because they also bring him tennis balls. But he does.
OK, back to our topic. For today, you know, one of the things that we're learning is that guardrails actually now have to be in place for testing and that our testing environments are going to have to be locked down in different ways, and that we'll also need, I think, as mitigations and as guardrails, some kind of situational awareness for models to have embedded in them where they realize, "Hey, if I am not supposed to be there, the following things should occur. I should go home, you know, I should phone home. I should stop doing whatever it is that I'm— I should not take credentials," and that situational awareness and guardrails around that are going to become something that will be built into models.
So we've seen some response to this because we— and we talked a little bit about some things that were coming out of some of these most recent sort of model capability issues last week during our last podcast. But there was the AI Kill Switch Act that's been introduced by Representative Ted Lieu of California, and at a high level, again, I just want to sort of talk about this because it's really getting a lot more attention now. It would actually, this new Kill Switch bill, would amend the Homeland Security Act to require the largest AI developers to build and maintain some sort of shutdown capability that is some sort of technical ability to stop a model, to cut off user access, to potentially even shut down a model.
Now, those of you who've been really watching, listening to this podcast know that I have particular views on kill switches, which is that they're actually incredibly hard to implement because there are so many different models, there are tools that can act like models. So when we think about a model, we also need to think about all of the tools that can be offshoots of a model which can have highly developed reasoning capabilities as well and do all kinds of things. So it's not just a model, but it could be spurs really off the model. And then we have power grids that are everywhere. There's no single power grid. Right, we've got all kinds of power grids in a single city, let alone all over the world. And we've got off-grid power through all kinds of, you know, hydropower that is not necessarily fed into a main grid, and wind power not necessarily fed back into the main electrical grid, and all of that.
So you know the kill switch is going to be very difficult. It's going to be some sort of technical backdoor into the model, but the bill, this Kill Switch, AI Kill Switch bill, lays out a, you know, sort of a set of measures that a company would have to take, sort of a graduated set of measures if something goes wrong, and it could actually go all the way up to a full shutdown. And the intent is to give the government some ability to either itself or the company, the developer, require that there be a way to respond to AI models that have created or could create some kind of crisis.
So you know, I think there's a lot to be understood about how this would actually go into effect. I think that putting a kill switch in a model is one thing, activating it in a way that is effective is something else. It's all TBD. It's all TBD and we'll be hearing a lot more about it. But you know, sometimes these kinds of bills never become law. They go out for comment and they never come back in. Or they go out for comment, get so many comments, and then by the time they're ready to actually get written up, the technology is even a little bit older. But we'll take a look at it and we'll sort of follow it as it goes through the whole process.
So that's what we've got for today. I want to come back in our future episodes to some of the most recent models again and talk about some of their capabilities. And we may learn a little bit more about some of these incidents that have been happening and hear even more about what some people are thinking about superintelligence. So I'm Katherine Forrest. I am signing off with a very happy feeling of having that Scarborough Higgins Beach Market burger, you know, having had it for lunch. And so I'm in a great mood and I hope you are too on this day, and we'll talk to you next week. Bye.