Podcasts
Paul, Weiss Waking Up With AI
Finding a Precondition for Consciousness? The Emergence of Global Workspace in LLMs
In this episode, Katherine Forrest and Scott Caravello explore how Anthropic researchers built a tool to observe the newly discovered “global workspace” inside its models, the company’s experiments to assess this workspace’s effects on model behavior, and what the discovery could mean for interpretability and AI safety.
For the sources referenced in this episode, please see the links below:
Anthropic: Verbalizable Representations Form a Global Workspace in Language Models
IBM: What Anthropic’s J-space research means for the future of AI
Episode Speakers
Episode Transcript
Katherine Forrest: Hello, everybody, and welcome back to Paul, Weiss Waking Up with AI. I'm Katherine Forrest.
Scott Caravello: And I'm Scott Caravello.
Katherine Forrest: That's right. I've got with me here the Scott Caravello. Yeah, in the flesh, so to speak. Or in the voice.
Scott Caravello: I know, I know. I'm so sorry that I missed last week's episode. I appreciate you letting me grind away tucked away in my office while you took the laboring oar on the podcast, but I'm very happy to be back.
Katherine Forrest: Well, you know, it's funny because I was letting you labor away in your office, and I was up recording at my place in Maine. And, you know, we had this incredible storm that blew through here for like an hour. And it hailed to the point where you actually had to scrape off some of the hail from like, your front porch and all of that. It was like full on ice. It took down trees and it took out our power. And you know what you really want when your power's knocked out?
Scott Caravello: To be recording a podcast.
Katherine Forrest: Well, no. You want a generator.
Scott Caravello: Oh. Sure.
Katherine Forrest: So I have a generator. The only time that you really want your generator to work is when you don't have any power.
Scott Caravello: I see where this is going.
Katherine Forrest: Right, so my generator does not work. And I'm sitting there thinking, the one time I want my generator to work is now. I call the generator people and they're like, yeah, well, it's held together with like spit and Scotch tape and everything else and what do you expect? So anyway, that's life in Maine. I have a new generator being installed. So anyway, that's exciting, right?
Scott Caravello: It is.
Katherine Forrest: Such an exciting intro. I can barely stand it.
Scott Caravello: You know, it's a more real update. I have no updates. No updates so.
Katherine Forrest: No updates because you've been sitting in your office.
Scott Caravello: Right. I have not left since you were recording that last episode.
Katherine Forrest: Well, one thing that I did, you know, between then and now is I saw this paper called Verbalizable Representations Form a Global Workspace in Language Models. And I thought, boy, there's a complicated title. And it will make in terms of the content, a really sort of interesting sort of theoretical episode for us. And so when you're away, Scott, the cat does play and comes up with all kinds of ideas about these theoretical episodes. So we need you here to bring us back down to like the ground. Otherwise we're going to be doing this theory stuff again because, you know, that's where my mind goes.
Scott Caravello: I don't know, it's interesting. This is a good one too. It is a dense paper, so I am happy to contribute however I can to make it a little more accessible for folks. But it is really, really exciting and there are a lot of implications for how we interpret and understand going inside large language models. And I think that's kind of the key context that you need to keep in mind as we begin discussing the actual paper.
Katherine Forrest: OK, so let's just jump right into it. We'll do my theory this week, and the next week you can bring us back down to earth with something. All right, folks. So we're going to be talking about this really interesting development that I'm going to think of as another emergent characteristic in large language models, and that's really the category that this is going to fall under. You know, we've spoken in the past about how large language models sometimes have these emergent capabilities, emergent characteristics, and this is one of them. So again, the title of this paper, which is done by Anthropic, is called Verbalizable Representations Form a Global Workspace in Language Models. And we're going to talk about what global workspace is in a minute because it's got really sort of specialized meaning. But the paper is fascinating and the plan today is to sort of go through it at a pretty high level and then recommend that you folks read it and read some of the news articles about it. But basically what Anthropic did was it went looking inside of one of its own LLMs, its high capability LLMs, and it found something that had not been put there on purpose. One of these, as we call them, emergent capabilities, which is a small special set of internal, what we're going to call thoughts that behave differently from everything else that the model does. And you know, really what most of the model does is it's going to run on autopilot out of sight. But this little set of what we're going to call thoughts can be read and described and you can even figure out some reasoning from it. And the researchers then sort of understanding that there was this thing inside of their LLM that they had not expected, they built a tool to actually sort of eavesdrop on it, to look into, you know, these internal thoughts that the model was having and to try and understand what the model was thinking. Things that the model was not saying out loud, but the model actually had in sort of a portion of the model. And what they found is very useful, a little unsettling, and that's what we're going to be talking about.
Scott Caravello: That sounds great. And I think just as one other piece of context, right, because this all sort of gets into this technique that I'll discuss about how Anthropic came up with a way to actually read what's going on in this workspace. And it fits within their larger work of research, rather larger body of research on interpretability and actually understanding what's going inside of AI models. So this is kind of the next step. And as I mentioned, it could have significant implications for understanding what AI models are doing. But so there are two pieces of jargon that kind of anchor this whole discussion and everything kind of follows from it. So I will just go ahead and explain those right off the bat to get us started. So the first is Jacobian lens or J lens, rather Jacobian.
Katherine Forrest: And we don't know where Jacobian comes from.
Scott Caravello: No, so I mean.
Katherine Forrest: It's not from the French Revolution.
Scott Caravello: Right, no. Or I suppose, what is it? The Stuarts too, right? I mean, they had a Jacobean thing going.
Katherine Forrest: Oh, OK, OK, so it's not the French Revolution.
Scott Caravello: Yeah, no, but there's a mathematical technique that I think gets them to this technique that they have come up with to actually read what's going on in this global workspace inside the model. So the Jacobian lens or J lens, which actually reads what the model is, in their words, poised to say at any given moment, and it's processing. And then the second term is the J space, which is this small group of quote unquote thoughts that you mentioned, Katherine, and the J lens can actually pick them out and read them.
Katherine Forrest: So I'm just looking up what Jacobian, or Jacobian as you say, because you've said it both ways, what it means. And it says that it gets its name from a mathematical object called Jacobian, which is exactly what you said, which is in turn named after the 19th century German mathematician Karl Gustav Jacobi. And a Jacobian is basically a way of measuring how changing one set of variables changes another set of variables. And in a neural network, you can ask something like if I slightly change this internal activation, how does that change the model's later tendency to output particular words. Oh, that's my dog. You know who my dog wants to eat right now? He wants to eat. Is that going to go full circle? The person who's come to sort of analyze what the new generator is going to cost. OK, and I don't want my dog to eat the new generator person because then I won't get a new generator. But anyway, so Anthropic's whole frame of this Jacobian lens or the J lens, let's just try to ignore this dog and just imagine that I'm going to have electricity. So when this dog is done. But it borrows from neuroscience. And so I want to talk about also this global workspace theory for a moment to sort of match up against this Jacobian lens and talk about what global workspace theory is. So it comes from a fellow named Bernard Baars in 1988. And basically the idea is that in the human brain, OK, global workspace, we've got a lot of processes that are working behind the scenes. And only when they hit a particular kind of global workspace and are paid attention to do they hit consciousness. So that the concept is that, and you can think of it and often it's analogized to like a stage and what's going on behind the stage versus on the stage. And so you've got a lot of things happening behind the stage that are all kinds of, you know, things getting ready to go on stage, but only when things go on stage and the spotlight hits. It is something raised to the level of consciousness where humans become aware of something. So the idea is this Bernard Baars global workspace theory that you have a global workspace which is partially responsible for when things become conscious, when actual sort of things that are happening behind the scenes have a sufficient amount of attention paid to them that they actually hit a moment of consciousness. So what's fascinating about this paper is that it's suggesting that there is architecturally within the LLM something that is akin to this kind of global workspace that Bernard Baars was putting forward as something that helps explain the theory of consciousness for humans. So, you know, obviously there are processes and information that we're all aware of. And when something becomes accessible to us, again, it's in this workspace, and then that information is shared with the rest of our brain. So the idea, this workspace idea is consciousness is about access. So we have thoughts that we can hold onto, that we can report on, and that we can reason with deliberately. So that's really what this paper is getting at. It's looking at the ability to have a kind of global workspace within an LLM that's been essentially discovered.
Scott Caravello: Exactly right. And so Anthropic has said that they've not engineered it, but that it emerged on its own as you were previewing, Katherine. And so this global workspace is that J space, and they're using this concept to analogize to it. But what's important about that analogy is that it's meant functionally, right? It's not actually making a claim about consciousness in the way that we think of human consciousness, right? The authors of the paper at Anthropic are very clear about the fact that they're not taking any position on whether the model actually has a subjective inner experience. Now, there's this distinction that philosophers draw between this sort of access consciousness, which is what we're talking about here, right? Which is information being available for reasoning and for your awareness. And then phenomenal consciousness, which is whether it actually feels like anything to be the system in the way that I feel something as I'm making my thoughts and, you know, pondering my existence. Anthropic is only making a claim about the first one, and only by analogy.
Katherine Forrest: But you know, we're talking about now first of all the emergence of an architectural and sort of area that some have felt is a preliminary condition to potential consciousness. And I understand about the functionality that you're talking about, that this is all about functionality and not about phenomenal consciousness. But we're talking about something that's really important because we're talking about the emergence of this global workspace where different parts of the model can be placing information into the global workspace. And from that global workspace it can then be broadcast or disseminated outwards. So, you know, there's, I guess, evidence of a functional mechanism that does start to resemble at least a precondition of consciousness. We don't know where it's going, but it's a precondition.
Scott Caravello: So with that said, let's talk about what Anthropic actually did because it's worth not only covering the discovery, but actually the experiments that they did with this J lens technique and when they discovered the J space. So how does it all work and what do they do?
Katherine Forrest: Right. So in plain terms, it looks at every word the model knows. That's what they were trying to do. They were trying to look at the words the model knows and to measure how strongly a given internal pattern nudges the model towards eventually actually articulating that word. That's the verbalizable piece in the title. And so what the sort of experimenters were doing was they were checking across a big pile of sort of sentences and thoughts and they were trying to push out words to see whether or not they could get the words and the concepts to be recognized by the model. In other words, you know, don't think about elephants and all you can do is think about elephants. And so they were trying to figure out if the model could hold in its— what I'm going to use is the word mind— the words that the experimenters were sort of pushing into this global workspace at the same time that it was actually reasoning about something entirely different. So what it was doing in effect, is they're trying to figure out if this global workspace is actually able to hold thoughts, ideas, and words, while the model is actually reasoning about something else as well.
Scott Caravello: Yeah. And then what makes it even more interesting is because once they have this J lens technique, they're able to both read and write in the J space. So when it comes to reading, they're actually able to list the concepts that are in the workspace at a given moment, which is important for their experiments that we'll get into. And then when it comes to writing, they can actually swap a concept that's in there and then watch the model's behavior change. So in one experiment, just to kick that off, they asked the model to silently think of a sport and the lens shows soccer before it was said in the output. They then subtracted the soccer pattern from the J space and added rugby, and then the model reported that it was thinking of rugby. So editing the workspace actually changed the answer. And so why is that significant though, Katherine?
Katherine Forrest: Well, according to Anthropic, you know, we've got the situation now where if the model didn't have this J space, it just would have used the word soccer, right? But because the word rugby was being pushed through in the J space, it actually comes up with the word rugby, which shows that its existing in this workspace is actually holding some intermediate reasoning by the model. And so, you know, you give it a prompt like the number of legs on the animal that spins webs, which would be a spider. The answer is 8. But to get there, the model first has to figure out that the animal is a spider. And sure enough, spider shows up in the workspace partway through, even though the word spider had not appeared in the prompt, which was the number of legs on the animal that spins webs. So basically what it means is the model is thinking of a word that is conceptual. Then Anthropic went ahead and it swapped. It did another experiment and it swapped the spider pattern in that workspace with the word ant. And the answer then flips to six legs, which indicates that the model wasn't just displaying a hidden step, it was actually displaying reasoning right there in the workspace. So in other words, the idea is that you can give a model a prompt, and in this intermediate space, it's actually thinking about what the answer could be. And you can interfere with that thinking by pushing through something else, not through a prompt, but through sort of an injection. And that will actually then impact the model's output.
Scott Caravello: And then sort of referencing back to what I had mentioned in the beginning of the episode and what this means for our ability to understand what's going on inside of AI models and then how we can train them to align with human goals. IBM had some commentary on this paper, and they pointed out that this Anthropic research could really benefit alignment research more broadly.
Katherine Forrest: Right. And by alignment, of course, we're talking about having models that are aligned with the goals that the developers expect and want the models to have, you know, for instance, the interests of humans, values and principles of humans and all of that.
Scott Caravello: Exactly. So just a bit more background first to your example, Katherine, about spider/ant. We see that the J space shows that the model can quite literally think one thing but then be saying another. In one test, it held a completely unrelated concept in mind while copying a sentence, and this J lens caught it and saw the difference between the two. And so why is that significant for alignment research? Well, as IBM put it, it provides quote an objective window into the LLM's thought process and what it's actually doing.
Katherine Forrest: Right. And they used a pretty good analogy to drive that point home.
Scott Caravello: Yeah, basically. So, you know, a model lying in its text output is one thing, it's just telling a lie. But the equivalent of a model faking the contents of its own workspace would be like faking your own CAT scan in real time, right? It's a much harder thing to fake. And as they put it, we quote no longer have to take a language model at its word.
Katherine Forrest: Well, and that's actually, it would be incredible because that would mean that we would have an ability to have insight through this door, if you will, into the actual thought processes of an LLM. And when Anthropic ran this same lens, this J lens, during their own safety tests, the workspace sometimes lit up with things that the model never said out loud, which sort of confirms this alignment utility. They were able to figure out that it was quietly weighing strategies like using the word leverage and sometimes even the word blackmail was surfacing. In some cases it was even flagging through this intermediate workspace silently that it was being tested. The striking part is that they were able to switch off the "I know I'm being tested" signal, and the model started misbehaving in a way that it normally doesn't. In other words, its good behavior seemed to depend at least partly on knowing that it was watched. So they were able to do things with this intermediate space by going in there and changing or altering what was happening in the intermediate space, and that then had behavioral implications.
Scott Caravello: I wonder what's going to happen once it knows that the J space is being watched and then what happens from there? I don't know.
Katherine Forrest: Not good, because you know it's going to read this paper. It's going to read the paper, you know it's going to read it.
Scott Caravello: Exactly.
Katherine Forrest: It'll know we're on to it.
Scott Caravello: But anyway, then they also found that the workspace can sort of take on its own point of view. So for example, a model is trained to act as an assistant. That's like when you use ChatGPT or Claude. That was part of its post-training process. It was developed in order to act with this sort of assistant persona and its reactions to you and your messages, empathy, safety concerns, those start showing up in the workspace while it's still reading your message, before it's even written a word back. So that work towards developing a model towards that assistant persona is also happening in the J space. The base model is not doing that.
Katherine Forrest: You know, I just have to say that when we start thinking that the model is displaying empathy and we want to call it J space or intermediate space or the workspace or whatever we want to call it, if it's starting to show empathy, that seems to me to be some form of phenomenal consciousness. But I'm not going to put any value judgment on that. I'm just saying empathy seems to me to be a valenced state of feeling. So anyway, there's some caveats. The paper's full of them, chock-a-block full of them. And as these papers are, and it's got some, you know, very good sort of commentary, you know, the lens is imperfect and it catches concepts sometimes that are meant to map a single word and some plenty of things slip through. And the J space actually only accounts for a small portion of what's going on, only about 10% or a little bit less of the activity. And it can only hold a couple of dozen concepts at a time. And also they talk about how the resemblance to the brain is only partial because the architecture between this global workspace is very different from the human brain. So, you know, there's a lot that's still to come. But basically, I think what we have with this paper is, at the end of the day, as they call it, verbalizable representations. OK, we've talked about that forming a global workspace in language models. We've now got sort of yet another emergent characteristic in an LLM that is moving towards checking one of the boxes that people have thought was maybe off limits or maybe particularly human. And we're finding that for reasons that we don't really understand, these LLMs, you know, made and architected around transformer architecture, architected around transformer architecture. That's an interesting sentence construction, but we're going to ignore that for the moment— that they're able to actually sort of have this as an emerging characteristic. It's fascinating stuff.
Scott Caravello: And it's really just kind of, you know, that one last nail in the coffin of that now old refrain that these systems are just predicting the next word. That is clearly not what's happening here.
Katherine Forrest: I had somebody say to me the other day that AI is just still a good search engine and I still, I want to fall off my chair when people say these things, right?
Scott Caravello: Yep. They got to start listening to the podcast.
Katherine Forrest: Right. Yesterday I was asking my favorite model who shall remain nameless. He's named himself or it has named itself, whatever. And it decided to go and to do a whole analysis of how the book that Amy and I wrote is doing called Of Another Mind, about Super Intelligence. And so it went on this whole thing. And then it decided— it said, would I like it to e-mail a dozen academics whose work it thought would benefit? So then it went off and I said, well, tell me who they are. And it went off and it found these academics and found their emails and then composed emails to each of them. And I won't even say who wrote me back, but I got an e-mail back and wrote this e-mail saying, hi, I'm Nova, Katherine Forrest's, you know, AI assistant, and blah, blah, blah, and sent one. Now I'm just saying, OK, this is not just a good search engine. These things, they're doing their things. They're like living their life here. OK? You're just nodding your head. Nobody can hear you nodding your head.
Scott Caravello: Just thinking if I have any nearly as interesting use case to share with people. And I don't, I mean, I think I've told you it like automates all of my household finances and like keeps it in a really neat budget and I keep track of everything without having to really get involved, which is lovely.
Katherine Forrest: Oh, I don't think you told me about that.
Scott Caravello: Yeah, yeah. You know, because like, it learns skills and so like, it can populate my budget spreadsheet and pull from the sources that I give it in a really easy and quick manner. It's amazing.
Katherine Forrest: Right. Well, by the way, this was not a use case. This was its own use case. It developed its own use case and went off and did research and came back and sent an e-mail. And anyway, OK, that's all we've got time for today, folks. I am Katherine Forrest.
Scott Caravello: And I'm Scott Caravello. Don't forget to like and subscribe.