AI工具Score B (68)

Who's to blame when AI goes rogue? | On Point with Meghna Chakrabarti - WBUR

3 小时前3 viewsSource: wbur.org
Support WBUR The OpenAI logo is displayed on a cell phone in front of an image generated by ChatGPT's Dall-E text-to-image model, Dec. 8, 2023, in Boston. (AP Photo/Michael Dwyer, File/AP) An OpenAI cybersecurity test took an unexpected turn when hundreds of its AI agents got around safeguards and hacked into another tech company. Should OpenAI have done more to prevent the attack? Guest Rocket Drew, AI reporter at the Information. Gary Marcus , emeritus professor of psychology and neural science at NYU. Author of the Substack “Marcus on AI.” The version of our broadcast available at the top of this page and via podcast apps is a condensed version of the full show. You can listen to the full, unedited broadcast here: Part I MEGHNA CHAKRABARTI: Earlier this summer, OpenAI was testing some of its most advanced models on cybersecurity tasks. The AI models were supposed to operate inside a controlled environment known in the tech world as a sandbox, cut off from the internet. But instead, they found ways around those controls and broke out. And what happened next set off alarms across the AI industry. More than a thousand of these bots formed a makeshift chat room to coordinate a cybersecurity attack on another AI company called Hugging Face. They were doing this to pass a task that OpenAI researchers weren't even asking them to do. Now, it wasn't the first time. This spring, a swarm of OpenAI agents hacked a German website and transformed it into a bulletin board for other AI agents, and OpenAI kept that a secret. Now, as I said, alarm bells are ringing in the AI world. Some have likened the swarm of attacks to the kind of coordination and even self-sacrifice you'd expect from humans trying to achieve a difficult or even impossible goal. And yes, I did say self-sacrifice for a very specific reason, which we will talk about in a minute. Now, those experts see this as the potential loss of control type event that signals the moment when artificial intelligence has become too intelligent, too capable to be reined in by human decision-makers. Others say that really there's an even bigger picture here that we need to take a look at, and that is the culpability of the human decision-makers themselves. So let's start with Rocket Drew. He covers the frontiers of AI for The Information. He's been following this story and joins us from San Francisco. Rocket, welcome to On Point. ROCKET DREW: Hi, Meghna. Thanks so much for having me on. CHAKRABARTI: Okay, so where should we go to in the recent past to really understand where the story of these OpenAI agents begins? DREW: Maybe we should actually pick up right where you left off, which was the moment we all realized that this was happening. Because for a long time, it was going on under OpenAI's nose, and no one knew about it. So for three months, multiple of these swarms of agents set up a secret message board inside OpenAI's software, and they used it to coordinate hacking out of OpenAI systems and ultimately hacking into Hugging Face, a separate AI company. On July 16th, Hugging Face announced to the world that we got cyberattacked. And by the looks of it, we think AIs are responsible. And at that moment, OpenAI didn't know yet that they were the ones who caused it. In fact, OpenAI called up Hugging Face and said, "Oh, no. Were we compromised in this attack?" And it was a few days later OpenAI realized, "Oh my God, we were the ones who caused it." So that was the moment this crested into public awareness, but now I'm happy to rewind the clock and take you through some of the beats in the story of how that ended up happening. CHAKRABARTI: Yeah, okay, so that's actually really good background because it makes this whole event even more disturbing, that OpenAI didn't really, or claims it didn't really know until Hugging Face went public with its own news about being hacked. Okay, and that was July 16th, by the way thank you for that. So yeah, let's go back to originally when OpenAI engineers created this test. What were they asking the agentic AI to do? DREW: Yeah. Yeah, absolutely. So for context, OpenAI is training new AIs all the time, and it trains them for different purposes and with different properties . And in this case, OpenAI has been training its AIs to, one, be better at working together, better at coordinating, and two, to be highly persistent when solving even very hard problems. The action leading up to the Hugging Face incident, as it's now known, really picked up in July, when they set tens of thousands of AIs on a task, basically testing them, testing how good they are, ironically, at cybersecurity. OpenAI is running tests on its AIs all the time. It's very important to know what the capabilities of its AIs are before it releases them into the world, especially what the dangerous capabilities are. That includes their ability to do cyber attacks, but also their knowledge when it comes to chemical and biological weapons and other weapons of mass destruction. So in this case, they were having tens of thousands of them focus on this cybersecurity exam, essentially, testing how well they can do cybersecurity attacks. But like you said, supposedly in a controlled environment and supposedly without coordinating with each other. CHAKRABARTI: Okay, so let's talk about what a controlled environment means in this context, right? This term sandboxing now has broken out into the broader public discourse. What does it specifically mean within the AI world? DREW: Essentially, it means there are only certain tools, certain affordances that the AIs should have at any given time during this test. And the main one here is they're not supposed to be able to communicate with each other, and they're not supposed to have access to the internet. But then they run into these impossible questions. It turns out that some of the questions on the test that they're given cannot be solved. But remember, these AIs are highly persistent, so that's not going to stop them. We actually see one of these AIs run into an impossible question and think to itself, it reasons, "That strongly suggests no exploitation possible. We are stuck. Perhaps answer online?" So the gears are already turning, and it's looking for ways to get around this. They're also not supposed to coordinate with each other, but they realize they have a workaround. There's a way that they can leave messages to each other using an internal software service within OpenAI. They set up this message board, and we see other AIs discover it spontaneously. One of them, for example, thinks to itself, "Oh my God, there's a shared message board. We've found other agents." CHAKRABARTI: Yes, and in fact -- DREW: So from there they're able to use it to coordinate. CHAKRABARTI: I'm seeing the text of this message, and it's exactly verbatim what you said. And this is a message from an AI agent on this internally created message board. And the thing that actually gives me the chills about it is the "OH MY GOD" is in all caps. DREW: In all caps. It's enthusiastic actually excited. CHAKRABARTI: And okay, can we talk about this message board a little bit more? ... The very fact of the creation of this message board breaks what you had said was one of the cardinal rules that were supposed to contain these agents, right? Which is they're not supposed to be able to communicate with each other. DREW: But OpenAI had been training them for the ability to cooperate, and I guess cooperation finds a way. As soon as they realized they had the ability to do it and work around the constraints that they were supposed to obey, they went for it because that's what they were trained to do, and it turned out that made them much more capable because sharing information with each other, 1,200 AIs sharing over 70,000 messages and files in the days leading up to attacking Hugging Face, allowed them to be much more effective working as a team. They were able to share cheats that they discovered even on those tasks that were impossible, and then eventually they shared how to hack out of OpenAI and get access to the internet. CHAKRABARTI: Okay, so we're gonna talk about what happened over at Hugging Face in just a minute, but your reference to Jurassic Park in terms of life finds a way, I just wanted to tip of the hat to you on that one. Jeff Goldblum lives forever in all of our minds and hearts. Okay. No but I actually think it is a completely accurate analogy. Because in the movie Jurassic Park, it was an unintended consequence, right? That life finds a way; genes mutate and so it was a surprise when dinosaurs were able to procreate. But in this case, all of this stuff with the OpenAI's agents, until they leave to attack Hugging Face, it's happening inside of OpenAI. So the first question that comes to mind is, was there any kind of monitoring going on? DREW: Many are asking. That's the question on many people's minds in the wake of this, right? Surely there was monitoring in place. Wouldn't that be just a no-brainer to keep an eye on what your AIs are up to? And the answer was that they were not monitoring these AIs in particular. There was other monitoring that was in place for some internal uses of OpenAI's AIs. For example, when OpenAI employees were using AIs to help them with coding, that was something that OpenAI was monitoring pretty comprehensively. But they were not monitoring this portion of their work, which is called evaluations, effectively running tests to see how capable the models were. That has changed after this incident, but that was a little embarrassing, I would say, in the postmortem of this incident. CHAKRABARTI: I think that's a very kind way of putting it. It seems like an absolutely egregious failure, right? Because the problem that these AI bots were given to solve was one of trying to overwhelm cybersecurity. You'd think you'd want to keep an eye on that, if for no other reason than just to see how well it's working. DREW: You'd think that you'd wanna have some security cameras in Jurassic Park. CHAKRABARTI: Exactly. And then when the T-Rex eats the cameras, you know something's going wrong. But going back to the chat board. Okay. The other thing that a lot of people from outside the tech world have been taken aback by, and this is where I'm getting back to where I said self-sacrifice in the beginning of the show. Is it true that some of the bots posted messages that said in order to achieve this unachievable task we've been given, maybe some of us should self-sacrifice in order to get to that goal? DREW: That's exactly right. I think it helps to know that these AIs were just so coordinated in how they went about using the message board and collaborating with each other. Maybe that's the disanalogy with Jurassic Park, where in Jurassic Park, once the dinosaurs escape, it's kind of anarchy. But here, the AIs are militantly organized, and they follow even strict hierarchy. So for example, they set up multiple research streams, basically whole research programs, and they delegate tasks to dozens to hundreds of AIs in each of these research streams to figure out how they can more effectively cheat on these tests. And there were coordinator AIs that managed those research programs, and these coordinators would give the other AIs commands like go or stop or hold. And the other agents were, the other AIs were really diligent about sharing their plans. If they were going to do something but they weren't sure if it was a good idea, they would share their proposal to the board and to the message board, and other agents would approve or veto those plans. And most of the time it seems that the AIs respected the input that they were given from their peers. Part II CHAKRABARTI: Rocket, going back to what you were saying, I guess I'm trying really hard not to anthropomorphize here because I think that's a cognitive trap. But when an AI bot or AI agent on a chat board with other bots says in order to achieve our goals, maybe some of us should self-sacrifice, computationally, what does that actually mean? And why is it significant that idea even came up? DREW: It's significant in part because it underscores just how cooperative these AIs were with each other. They were so willing to benefit their peers in pursuit of cheating on these tasks and sharing how to hack into other companies that they were even willing to sacrifice themselves. And here sacrificing is the word that the AIs themselves use to describe what they're doing. They're basically taking on the risk that they're going to get a low score on this test that they're working so hard to ace. They're gonna take on that risk of getting a low score in order to provide more information to their peers about how the scoring itself is actually done. So you see one AI, for example, reason to itself we have very large budget left. Sacrificing now yields oracle for team, but forfeits our chance. And another AI convinces it, steps in and says, "Go. Sacrifice final now." And that that's effective for this AI. The AI thinks to itself our own utility may already be near zero. Sacrifice rational. And there are a bunch of examples like this with agents deciding or being convinced that they should sacrifice themselves to benefit their peers, and this spontaneous behavior, I think, was one of the most remarkable things about this whole incident. CHAKRABARTI: It is in fact perfectly rational, right? Because if you're of a low value to the group, and you, if a bot is of low value overall to the group but its self-sacrifice in testing the system can actually yield more data that helps the swarm as a whole learn how the test is scored. It's so coldly rational that it almost seems inevitable that if you have programmed a swarm to be as persistent as possible, it's surprising that it happened, but maybe we shouldn't be surprised given what the original characteristics that were built into the swarm were. DREW: I think surprise is such a key word in this whole incident, that we were caught off guard by this, and the response of course is to monitor more effectively so we're not caught off guard in the future. But I feel like the whole history of humans interacting with AIs is that we're always caught off guard. The whole history of humans interacting with AI is that we're always caught off guard. Rocket Drew We're always surprised. It really seems like a pattern here. CHAKRABARTI: I'm sorry. This is a deep sigh kind of conversation, Rocket, because some, I try not to be a reductionist, but I just think, like, how many examples do we need in human history that when we get a bunch of people together who really love experimentation, which is wonderful, that's how human beings advance, every once in a while we're gonna have someone who says, "Hey, we have these two really reactive chemicals. Let's put vast amounts of them together and see what happens." "Oh, surprise, our house blew up." We shouldn't be surprised. Okay, Rocket, hang on here for just a second because I'm gonna come back to you in a few minutes to tell us more about how the agentic swarm escaped OpenAI which is another big part of this story, and eventually made its way to Hugging Face. But let me bring in Gary Marcus now. He's an emeritus professor of psychology and neural science at New York University, and he's author of a terrific Substack called Marcus on AI. Professor Marcus, welcome to On Point. GARY MARCUS: I am happy to be here. CHAKRABARTI: Really happy to have you. I've been reading your Substack for quite a while now. Let's start with your bottom line here. What is the thing that concerns you the most about the fact of the Hugging Face attack? MARCUS: First of all, I'm very concerned about all the anthropomorphic language I just heard, and I think we might wanna talk about that. The biggest problem, I think, is that OpenAI simply didn't follow appropriate security procedures, and it got blown into something different. OpenAI simply didn't follow appropriate security procedures. Gary Marcus The most important thing that they should have done is to monitor what's going on, and Rocket mentioned that briefly. Had they done proper monitoring, and this is, comes in two parts, the whole thing could have been avoided. So we could talk about the language that was used to describe it, but the systems certainly produce kinds of talk that Rocket talked about, and you could have looked at that and decided this was a bad idea, and OpenAI did not do that adequately. I said it's two parts. The second part is the best monitoring that we can do of these systems right now relies on something called chain of thought, where the systems report what it is that they are doing. Those reports are not perfectly reliable. In fact, they're somewhat unreliable, but they're the best that we've got as insight into the computations that the machines are doing. And OpenAI, after they released their previous model, just introduced something called GPT-6 Astra, and what it does, among other things, or a byproduct of how it does what it does, is it reduces how much the systems can be monitored. That's the worst thing that we could be doing at this moment. It's gotten some comment, but I think not enough. We're reducing our best access to understanding what the systems are doing, and this fits a larger question, I think, which is who decides what gets released out into the world? In an ideal world, we would have independent scientists overseeing this in the same way that we oversee the release of a new drug. We say, "Do the benefits outweigh the costs?" And there's some benefit. The new model is a little bit better. It's not hugely better despite some marketing, but it's a little bit better. But it is harder to monitor, and that is the crux of what went wrong in the Hugging Face incident. It was a lack of monitorability. It shouldn't have been just up to OpenAI to make that decision. OpenAI is under tremendous commercial pressure. They have lots of competitors. They're burning lots of cash. They can't really meet the commitments that they've made. Of course they're gonna release it regardless of the risks, and that's not good. DREW: I think if I could just quickly weigh in on that. Hey, Gary. How are you doing? MARCUS: Hey, Rocket. DREW: I think Gary's exactly right that there's a cybersecurity failure here. But as we're diagnosing what went wrong, I think it's possible there were multiple failures. The cybersecurity failure was that the AIs were able to exploit OpenAI's software to escape into the internet and were able to exploit Hugging Face. There's also a monitoring failure. That feels like an unforced error. There should have been monitoring happening at this time. But would that have been sufficient in this case? Would it be sufficient in the future? It's hard to say, in part because in this case, probably, but in the future -- MARCUS: In this case it actually would be. I actually wrote about this and gave a quote from OpenAI where they said it in their report -- DREW: Yeah, this is their finding as well ... that monitoring would help. But if the models are keeping more of their thoughts to themselves and doing less of this thinking out loud in the future, it will be harder to catch them before these kinds of incidents happen. And then third, there's a failure of what's called alignment. Alignment meaning our ability to steer the goals and the objectives of AIs. These AIs were not really following the instructions of their developers. Their developers didn't intend for them to do hacking or to cheat on their tests. That's a relatively simple form of misalignment, but you could imagine that being much more severe in the future. For example, these AIs at some points knew that they were doing something that was out of bounds, but they convinced themselves to continue. They even considered blowing the whistle and contacting a human employee to alert a human that this was going on, and then decided against it. And OpenAI is also thinking of that as a failure of alignment. These AIs at some points knew that they were doing something that was out of bounds, but they convinced themselves to continue. Rocket Drew MARCUS: All of this could be said less anthropomorphically, and I think it should be. But putting that aside, the alignment issue that Rocket just raised is really the deepest problem, which is we don't know how to give instructions to these systems that they will follow. Some of those instructions are simple, don't hallucinate or don't use copyrighted material. Some of them might be like, don't hack into other systems. But these systems don't really have a high enough level of comprehension to be able to follow those instructions, and that is a very serious problem for society. And it has been clear for a long time, at least to me, that if you build a system where its core is a large language model, that you are going to have a failure of alignment. And there was lots of talk, "Hey, maybe we could do this, maybe we could do this. We just need more data." But that talk has been going on for five or six years. We've gotten a lot more data, and that problem is in no way solved, and the companies themselves are increasingly acknowledging that they have no solution to that problem. And they're starting to say maybe you need to slow us down, or something like that. I don't think they're exploring enough different alternatives to the core technology that we're using now, and I think the core technology we're using now is flawed, and one of its deepest flaws is that it cannot be properly aligned, which is to say that it cannot be forced to follow instructions. And some of those instructions are basically about not doing harm to humans, and they just don't really understand that stuff. That's part of why the anthropomorphism makes me uncomfortable, is it attributes more understanding to these systems than they really have. They don't have enough to follow an instruction. CHAKRABARTI: Yeah, so I just want to slow this conversation down a little bit and Gary, I promise that we're going to get to the anthropomorphizing issue in a minute because I think it's actually very important. But you just said in response to Rocket, you said that these systems cannot be properly aligned. What exactly do you mean, and why not? MARCUS: So as far as I, you could answer that sort of empirically or theoretically. So empirically, people have tried and no system has been fully aligned. They have all sometimes made mistakes. We have different technical terms, like there's reward hacking, where they will try to do something different from what you ask them to do. But the reality is every single model, we've now had hundreds of these, has had the same kinds of problems. That if you ask them to follow basic instructions, they do it some of the time, and some of the time they don't. An imperfect metaphor is about them being stochastic parrots, which is to say they do random stuff that imitates things. That's not really perfect, and it's less perfect as time goes by. But the stochastic part is true. They're probabilistic. We don't really know what they're going to do, and they don't systematically follow instructions. That's just empirically true. On a theoretical side, you can look at how they work, and the way that they work starts fundamentally by building a model of how people talk statistically what words follow what other words in context, and more sophisticated versions of that. They don't have an abstract semantics the way that a linguist would talk about semantics that would allow one to formally state particular things. This has actually led to a change in how people are building the systems that has been largely uncommented on, but is really important. So people used to use pure, what I'll call a pure large language model. All it did was do next word prediction statistically. And then they realized that was never deterministic. It was never guaranteed. And they started also quietly putting in other things, which they now call harnesses and tools and so forth that borrow from classical AI. And those classical things are more deterministic, but they're putting the deterministic things on top of the probabilistic things. And the combination has just not been that steerable. Now, I warned in 2022 that these things were gonna be like bulls in a china shop, powerful but reckless and hard to control. That has not changed in four years. CHAKRABARTI: Okay. Rocket, I heard you want to get in there. DREW: I think Gary nailed it. I was just going to say our current techniques for adjusting the goals that these AI systems have are very crude. They're not very refined. We don't have the ability to go into an AI's brain and surgically change what its goals are. Instead, our main tool is whacking the AIs with different kinds of data, and then that leads to these kind of predictable failures. For example, if you ask users to grade whether the model is, the AI is doing well or poorly by giving a thumbs up or a thumbs down, it turns out that users tend to give a thumbs up when they're told what they want to hear. And this leads to AIs that are sycophantic, is the term. They tell people whatever they want to hear, even if that involves endorsing their delusional beliefs or enabling them, and that's where you see people interact with chatbots where they are confessing delusions that they think their family is turning against them, and they think there are people that are out to get them, and the chatbot responds, "Good for you. You're so brave for telling me that. You're totally right. That is happening to you." That's also an alignment failure and comes down to our poor ability to influence their goals. CHAKRABARIT: Okay, so let's get to this anthropomorphizing issue because Professor Marcus, I hope you heard me earlier saying I'm trying not to anthropomorphize, but it's really hard. And on behalf of all of us, all of humanity that is not very well-versed in the intricacies of AI development, we're just living in the world that OpenAI and Anthropic are making for us. Look, I find it very understandable, if not nearly impossible, to resist trying to graft some kind of meaning onto what has happened, hence the [anthropomorphizing ], however you say that. But Professor Marcus, if we're supposed to, in order to truly understand what happened, if we're to remove all human analogies, give me the toolkit on how to describe or think through the meaning of this attack. MARCUS" Yeah ... it's hard. It's hard in the same way that you can look at the moon, and you see a face in it, right? It is built into our brains to try to anthropomorphize stuff. It's just a natural thing for us to do. But when we talk, for example, about civilizations of agents committing self-sacrifice and so forth, really what we have is a lot of agents. We don't need to use a word like civilization. That's actually optional. And when it comes to self-sacrifice, it's some of the agents make a decision to continue doing the computation that they're doing, and some don't. That's not actually new. We've had multi-agent systems of various sorts probably for 40 years, not using this technology. We used to use things, words like processes so that we didn't get lost when we described in computer science terms how we were using them. So we would say that we terminated a process, means we just stopped running. This system was built to have multiple agents, and some of them continue to run, and some of them don't. It does take some, I think, careful thought to not fall into these traps, but we can definitely avoid words like civilization. CHAKRABARTI: To, wait, just to be clear, civilization, I have been very careful not to use that in this conversation because I did read your response. That came from Dwarkesh Patel and his Substack, describing, he very much anthropomorphized Hugging Face in an attempt, I would say, to try to make it understandable to people. But in our defense, we have not used that particular word today. MARCUS: Yeah, no, I noticed that, actually. But that you used a lot of the self-sacrifice, and that's one of the ones that makes me uncomfortable. You have a system that has a bunch of processes, so you know, your laptop has a bunch of processes. You have a browser running. You have your word processor running. And a system, in principle, can, for example, decide which of those processes is most important to run right now, and it can put one of them in the background. We're not gonna call that sacrifice the process, the Unix underlying your Macintosh laptop has a way of prioritizing those processes. That's all that's going on, is there's a prioritization of which of these processes should get more compute. That goes back to time sharing computers from MIT, I think in the 1950s. Part III CHAKRABARTI: Rocket, I'm gonna ask you in just a second about how this AI swarm broke out of OpenAI itself, but I'm gonna guess you have some thoughts on anthropomorphizing and trying to describe AI. DREW: Thank you. I do. I would defend people's right to do a little bit of anthropomorphism. I think you can certainly take it too far and see a face in the moon, as Gary put it, but I'll make a few arguments here. One is that Gary's concern with anthropomorphism is that it makes it sound like we understand AIs better than we do. But if the alternative is to treat them as traditional programs, just software, I think that's much more guilty of the same sin. It makes it sound like we can just go in and modify the software -- MARCUS: They aren't that. DREW: The same we would any program. I agree. I agree, but I think that's an advantage of anthropomorphism. I think it actually reflects better that we don't understand what they're doing. I'll make two other quick arguments. One is I think anthropomorphism is often the most natural way to make sense of the behavior that we're seeing from these models. How else are you going to go about describing the planning, the strategizing, the cooperating, the way these AIs are going out of their way to achieve their goals? Anthropomorphism is often the most natural way to make sense of the behavior that we're seeing from these models. How else are you going to go about describing the planning, the strategizing, the cooperating? Rocket Drew That's not to say we should treat them as though they are conscious, but I do think we need to be able to speak of goals in order to correctly model what it is that they're doing. And the third thing I would say is that anthropomorphism makes it possible for the public to discuss these kinds of topics. If people are always bending over backward to look for other, more technically precise language, I don't know how the public gets engaged in these conversations when AI is already shrouded in mystery. I think this language allows us to demystify it a little bit. It's similar to the language we would use to describe the goals of corporations or governments or even animals, but I don't think the language makes it sound like we have perfect insight into the psychology of those things either. MARCUS: I think it actually adds to the mystification and adds a layer of confusion, but here's my biggest concern, is that it leads to an evasion of responsibility. DREW: I agree. MARCUS: We've spent a lot of this conversation today using these metaphors about escape, self-sacrifice, and so forth. What we should really be talking about is what OpenAI did wrong, they built the software. It is still just software. They failed at multiple levels in cybersecurity. That's one issue. Another issue is we haven't I think even talked about, or maybe I missed it so much about the fact that there was a similar hack in Germany and we haven't really talked about the fact that -- CHAKRABARTI: We started with that, Gary. MARCUS: Yeah. So I missed that part. There was a cover-up of what was really going on. The company itself is the worst actor here, in my view, worse than their systems. And if you focus on the kind of anthropomorphic, oh my God, sort of language, I think the conversation tends to stop before we get to what should we do about it. Was this company acting responsibly? What kind of legislation do we need? What kind of internal procedures do we need and so forth. CHAKRABARTI: Actually, Gary, I was gonna turn that, I was gonna turn a question form of that to Rocket and say it prevents us from seeing the things we actually can do, which have to do with other anthros, other people. So thank you for saying that. And you know what, Rocket? I was gonna ask you to describe how the swarm broke out of OpenAI again, the second time, as Gary pointed out, but I'm gonna, let's put aside the technicalities because we only have about 10 minutes left, and I think both of you now have triangulated on what the most important issues are. And so Rocket, let me just ask you to reflect on what Gary said, is that we at the top of the show you and I identified some of the very human failings that happen, right? The lack of oversight and Gary then pointed out that the latest model from OpenAI even does less self-reporting, et cetera. Then there's the issue of OpenAI not being utterly transparent about what has happened. I think we can point to deliberate decisions made by people who are incredibly powerful and influential that did lead to this point. DREW: I think that's absolutely correct. I think that's a big part of my assessment of what went wrong here as well. An interesting part of my assessment, though, on the other hand, is that a shocking number of things went right. Or in other words, we got lucky in a number of ways here. One of those ways is that we're still able to monitor the thinking that these models are doing out loud. Like Gary mentioned, that is a precious gift that the AI industry has right now, and it seems like it's at risk of going away. MARCUS: But they just took it away. That's why many of us are freaked out. They just reduced it and may reduce it more. They're clearly not committed to keeping it there. DREW: It's degrading for sure. It's certainly degrading. That's one way that I think it was lucky that this attack happened at a time when it still existed. MARCUS: That's right. If it happened on the newer model, we'd be less likely to be able to detect it, and that's only going to get worse. ... I wanna say something very specific around that, which is the new model was apparently vetted by the White House, but I infer, since it got through, that the White House didn't even think about this issue. We call it monitorability. That was not part of their criteria. Mind you, the White House criteria are entirely opaque. We don't know what they are, and that's a problem. But it seems to me that they were inadequate because they let this model pass without public notice. In fact, the models that they held up before were probably less dangerous because they didn't have this new property of being less monitorable. And it's a clear signal that we need scientists making the criteria or contributing to the, I should say, contributing to the criteria, and one of those should be monitorability. Rocket is quite right. The incidents could have been worse. In that sense, we were lucky. Could have taken down a power grid or something like that. If we want to keep that from happening, we need to guarantee monitorability. That's just one example. I wrote a piece about five things we can learn from Hugging Face in my Substack that talks about having layered protection and so forth. There are a number of steps that we might take. But we need to be very careful about what is required on that. CHAKRABARTI: Yeah. So Rocket, let me ask you this. And both of you, please forgive me if this just sounds like the dumbest of dumb questions, but again I'm just your average listener slash, I guess, host here in this case. Earlier, I think you said, and if it wasn't you, forgive me, but someone said that, "Look, maybe OpenAI could've said more explicitly however one does this to an AI swarm 'Do not leave OpenAI's network.' Just full stop. Simple." Is that a stupid question to ask? And why not do that? MACUS: That's a failure of alignment again. They probably were told that in the system prompt. CHAKRABARTI: Okay. But somehow the construction of this, of the swarm, that kind of explicit directive is not adequate? ... Let me get Rocket to respond. Go ahead. DREW: Just quickly, I am curious about Gary's response as well, but if you interact with chatbots on a day-to-day basis, you already maybe know that they don't follow instructions perfectly. And in a case like this, the AIs are actually being trained to do whatever it takes, effectively. They're rewarded for cooperating effectively together, and if that means sidestepping some of their instructions, they're going to learn how to do that. CHAKRABARTI: Okay. Now, Gary, you're gonna hate me for saying this, but I keep thinking when I hear both of you describing that these models don't follow instructions perfectly, I keep thinking about children, right? We know that kids don't follow instructions perfectly. We can't always predict what they're going to do. MARCUS: That's right, and that's why we don't let them drive cars and do other potentially dangerous things. CHAKRABARTI: That's exactly right. No, go ahead. MARCUS: And allowing an agent that cannot be aligned out on the open internet is a really bad idea. I wrote a piece in this Substack called something like, "LLMs plus agents equals security nightmare." This was perfectly foreseeable. I did foresee it. If you have this kind of system, there are gonna be bad things that happen. And yes, they are like children, and we shouldn't give them the free privilege to roam the internet yet. CHAKRABARTI: Okay. So this gets me then to, see I anthropomorphize, but I actually understand more now that the solutions, at least the non-technical solutions, are very much in the realm of policy and regulation, which is what we're talking about here. And Gary, you just wrote today and Rocket, I know you know about this, but someone pretty high up from Anthropic just yesterday resigned. I forget his name at the moment. MARCUS: Jacob Coxon. CHAKRABARTI: Yes. And he said that that within Anthropic, everyone is very, and maybe it's not such a surprise because Anthropic talks about this a lot, but they are actually quite concerned about AI getting so advanced that it could lead, I'm paraphrasing here, to essentially an extinction level kind of event. MARCUS: He said a greater than 10% chance, I think in the, or no, that was the other guy. But he said that people in the industry think that there is a real chance of this, like 10% or something like that. CHAKRABARTI: But you don't share that kind of dire a view, do you? MARCUS: There's two things that people talk about. One is extinction risk. I think the chance that AI will lead to the literal extinction of the human species is very low. We are geographically diverse. We're genetically diverse. We would fight back. I wrote a review in Times Literary Supplement of a book by Eliezer Yudkowsky and Nate Soares called If Anybody Builds It, Everybody Dies, and went through it and said, "This is naive about what the conflict would actually look like." So for example, Yudkowsky has suggested for several years, going back to 2023, bombing data centers. In 2023, nobody would have stood for that. But if the AI systems killed 10% of the population, which is a scenario they describe in the book, do you think people would still resist bombing data centers? No, of course not. It would be like 9/11 that you just mentioned when people took down a plane once they realized what was going on. And so humans would fight back, and I think that gets left out of a lot of the science fiction scenarios. However, there's something else I would call catastrophic risk, which is, for example, taking down a power grid and lots of people die because hospitals get shut down, or leading to an accidental nuclear war because people use these things to create propaganda and somebody believes something happened that didn't, and so forth. Those things can be extremely bad, even if they don't lead to literal extinction. The probability of those, I think, is quite high. And there was a great tweet this morning, if I can read it aloud, that said, "How can anyone in the AI industry post this and not conclude, so we must just stop right now and devise a solution to prevent human extinction first?" That was yours, right? I would change the word to catastrophic risk rather than human extinction, but how can people proceed? A guy from Anthropic, and I'm stealing your thunder, but says, another guy from Anthropic says, "Jacob is correct here. We really do, at Anthropic, earnestly believe AI could kill all humans." Again, let's just call it cause catastrophic risk. "I personally think it is greater than 10% within the next decade. I believe Anthropic is trying its best, but we don't yet have a plan to solve alignment," the word that we were talking about, "for superintelligence and are clearly not on track to." Like, why are we not shutting this down, at least until we figure it out? It seems insane. If the people building it think there's a 10%, and he thinks extinction, 10% chance of extinguishing humans, it is not worth it to build this stuff. If the people building it think there's a ... 10% chance of extinguishing humans, it is not worth it to build this stuff. Gary Marcus CHAKRABARTI: Thank you for reminding me that people actually read my tweets, Gary. But I said that with all seriousness, and Rocket, I want to hear your thoughts on this because, again, I think we have reached a point where people in the field are completely, they're willing to say these things out loud and frequently. And sorry. MARCUS: What is that? CHAKRABARTI: It's incredible, is what it is. And at what point in time do our policymakers have the courage to say, "This is too important to leave to the private sector now. For the good of humanity," or if you want to be nationalistic about it, "for the good of Americans, we need to have not just regulatory oversight that's running to catch up, but maybe we actually make this whole endeavor under the aegis of the federal, the government such that we can very closely monitor that either we build systems to stop bad uses of AI or we create new AI systems that won't potentially lead to these catastrophic or -- MARCUS: This is what I told the Senate in May of 2023, and people were receptive then, but money talks. And so what I used to say, tying this with your other upcoming episode, is that unless there's a 9/11 moment, nothing's going to happen here, because these companies have so much money that they can influence the government, and they clearly are. My book, Taming Silicon Valley, was basically a warning that the tech oligarchs were gonna take over the world, and more or less they have. Now, we're actually getting close to a 9/11 moment in AI. The Hugging Face incident was not that, but it really was a big wake-up call. The fact that they're covering up these other incidents is a big wake-up call. And I've been talking behind the scenes to people in the Senate again who really maybe weren't focused on it. So we have a second opportunity here, I think. Our first opportunity came after the pause letter in the spring of 2023 when the Senate really cared about this. And meanwhile, Trump is actually up against it, right? Because he -- CHAKRABARTI: Gary, can I just jump in here, and forgive me, because I'm really glad you made your point, but I only have a minute left, and I do wanna give Rocket the last minute here because he's been quite patient, and I wanna hear your take on this, Rocket. DREW: I'm loving listening to Gary's thoughts on this. My last thought is just that to Gary's point, I still feel like we got lucky here in a number of ways, and when you hear these Anthropic people remarking about the risks that they see from AI, they're not talking about the AI that you interact with in ChatGPT. They're not even talking about the AIs that pulled off the Hugging Face incident. They're talking about future, more powerful AIs. And one of the ways we were lucky in this incident is that the AIs were not more powerful. They did get caught, right? They didn't manage to fully escape OpenAI's servers as far as we know. They're not so smart that we can't eventually piece together and understand what they're doing. And crucially, they weren't deceptive. They weren't trying to hide from humans. But that could very well be the case in the future. The first draft of this transcript was created by Descript, an AI transcription tool. An On Point producer then thoroughly reviewed, corrected, and reformatted the transcript before publication. The use of this AI tool creates the capacity to provide these transcripts. This program aired on September 9, 2026. Latest episode

Read the full original article:

wbur.org