AI工具Score B (53)

How, Exactly, Could A.I. Kill Us? | The New Yorker

2 小时前2 viewsSource: newyorker.com
Save this story Save this story Save this story Save this story Listen and subscribe: Apple | Spotify | Wherever You Listen Sign up to receive our twice-weekly News & Politics newsletter. When Jacob Coxon, a mathematician and software engineer, resigned from his research job at Anthropic last week, he warned, “The people building AI earnestly believe that it could kill us all by the end of the decade.” This would be a remarkable statement were it not for the fact that artificial-intelligence leaders have long been saying precisely this. In 2018, the Anthropic C.E.O. Dario Amodei, then a research scientist at OpenAI, raised the concern that a superintelligence “could destroy humanity,” adding, “I can’t see any reason and principle why that couldn’t happen.” Earlier, in 2015, Sam Altman, just before he co-founded OpenAI, said, “I think A.I. will probably most likely lead to the end of the world, but in the meantime, there’ll be great companies created with serious machine learning.” Elon Musk, in 2014: “I think we should be very careful about artificial intelligence. If I were to guess at what our biggest existential threat is, it’s probably that.” Perhaps the only thing that’s changed between then and now is that the rest of the world is finally paying attention. In recent weeks, the same technology that, a couple of years ago, couldn’t count the number of “R”s in the word “strawberry”—and, a couple of days ago, insisted to me that Dolly Parton is still alive—has been used to solve the Navier-Stokes problem, which has been stumping mathematicians for nearly a century, and has also demonstrated its ability to go rogue in a series of disturbing hacking incidents. Last week, Anthropic also published a report detailing various ways in which bad actors have attempted to use the company’s A.I. models, including one especially troubling case of a scientist using Claude to study a virus at a military research institute—work that could yield a vaccine, a biological weapon, or both. A few days later, Amodei published a letter calling for an industry-wide slowdown and more government regulation, to which President Donald Trump responded, on Truth Social, “The only control or ‘guardrails’ that AI needs is a STRONG AND SMART (High IQ!) PRESIDENT, and the U.S.A. has that, in spades!” I recently spoke on The Political Scene podcast with my colleague Joshua Rothman, a staff writer who has been covering A.I. for years, about whether we’re all doomed, and what it would even look like for A.I. to destroy humanity. Can A.I. leaders save us from their own creation, and how can the government coöperate in order to do so? And is A.I.’s capacity to do good—its potential to mitigate climate change or innovate medical treatments—hopelessly intertwined with its capacity to do bad? Our conversation has been edited for length and clarity. A lot of people in the world of artificial intelligence are talking about their P(doom) number, which is the probability that artificial intelligence will lead to an absolutely catastrophic situation—possibly, or probably, killing us all. What would you say is your P(doom) number? My P(doom) is pretty low. It’s, like, ten per cent. O.K., so the same number that we’ve seen a lot of people in the A.I. industry use recently, right? Yeah. And I have to say, also, like, what does that even mean? It’s kind of like a vibe check on my disposition. It’s not like off-camera I have huge whiteboards covered with calculations. Well, if you said ninety, I’d probably end the interview right now and go, uh, do something about it. . . . Right. Well, it’s, like, ten per cent, and it has partly to do with the fact that there are lots of other things in the world that could cause really bad stuff. There’s climate change, and nuclear weapons, and bioweapons, and they each should have a P(doom) also. So you have to kind of try to keep things in proportion, as far as being terrified. But my A.I. P(doom) is probably ten per cent, which to me feels incredibly scary—like, super high. Like, way too high for my comfort. Yeah, ten per cent is still ten per cent more than ideally it would be. I think it’s fair to say that these really deep anxieties about A.I. entered the mainstream last week when an A.I. researcher and mathematician named Jacob Coxon quit Anthropic. But, at the same time, it’s not really a big secret that people in the A.I. world think that A.I. might kill us. Why do you think people are suddenly paying attention to this problem? It’s a really good question. I think there are sort of two sides to that. Like, one side is: Why are people in the industry sort of getting behind this at this moment? As you say, they’ve always been talking about it, which is important to remember. This isn’t something that they’re just bringing up now. A lot of other researchers have published essays or long posts on social media, unburdening themselves of their concerns as well. So the industry’s gotten behind it. And then people are paying attention, which is overdue. I think people are paying attention not because A.I.’s gotten better in their everyday lives but because of two things. One, these autonomous-agent hacks. Like the hack that occurred at Hugging Face—which is a big A.I. company—that was conducted by these A.I. agents that escaped from OpenAI’s servers, and, in an effort to cheat on a test, they hacked this other company on their own. And then there’s also this startling progress that A.I. has been making in solving math problems that no normal people understand, but that even mathematicians who do understand the problems find to be important. An A.I. model recently solved a Millennium Prize math problem, which is an actual breakthrough. And those two things together are what I think makes it suddenly seem real. In other words, on the one hand, the A.I.s are out of control. They’re acting on their own. They’re not trustworthy. They’re covering their tracks. They’re working together. They’re using weirdly emotive language to describe their own decision-making. They just seem crazy and unpredictable. And then at the same time, the models are really smart, and they’re doing things that—I don’t like the word “superintelligence,” and I don’t like the term “A.G.I.”—but they’re doing things that, like, I can’t do, and that most of us can’t do. And that combination of things together makes this discourse on the dangers of A.I., which has often sounded science-fictional and maybe even like it’s marketing hype or something, sound, like, really plausible and salient. I want to go back to the initial warning, which is this idea that there’s some chance that this could all happen in a decade, which is a span of time that is simultaneously very close and very far away. One of the lessons of climate change is that people can accept a threat intellectually, but still not act on it because it doesn’t feel present enough. Like, New York might be underwater in 2036. It’s really hard to imagine where I will be or who I will be in 2036. Do you think that we risk the same problem with A.I., where it’s this abstract threat that’s close and yet far away, and so we don’t actually end up doing anything about it because we can’t even really picture what that threat would even look like? I think it’s really helpful to disentangle two worries that are often conflated. And they’re both equally worrisome, so it doesn’t make it less scary to disentangle them. The first is about what’s sometimes called A.I. takeover. The idea is that the A.I.s will get super smart. They’ll decide, for reasons that make sense to them, in their bizarre way, to do things that we don’t want them to do, and those things might have really negative consequences for us. And people do worry about that. They worry about A.I. systems that are, like, really good at hacking, and they want their company to win or their country to win, and they take drastic steps or dangerous steps that no person would want them to take. Those are Skynet-type worries. That has to do with an idea often called superintelligence, which is that the A.I.s will just quickly get really, really smart, and then they’ll just, like, have no use for us any longer. And that is something I think is worth worrying about. And, like, Bernie Sanders and Greg Casar have legislation in Congress to ban the pursuit of superintelligence. But then there’s just a more ordinary way in which A.I. is dangerous right now—like, it’s already worrisome. Anthropic published a report in September that goes through some of the ways that people are using A.I., and they try to stop them from using it in these ways. It includes groups in Yemen who are trying to vibe code software for guided missiles. It includes people launching cyberattacks using autonomous bots, broadly similar to the ones that conducted the Hugging Face hack. So that’s not a distant danger that is abstract. It’s actually that right now, the technology as it exists can be misused by people or it can end up doing things that people don’t want it to do. That’s the technology that exists today and that sort of came into existence relatively recently. But, of course, it is always improving. So it becomes a question of: If we just take one step of improvement forward, do we reach a point where it becomes very, very difficult to control it in the here and now—totally separate from those larger, more abstract sci-fi scenarios, which we also need to be concerned about? You just laid out two different scenarios there—A.I. becoming a godlike entity that decides to go rogue and kill us for reasons that we might not even understand, and humans misusing these powerful tools to create something that could harm other humans. Is it right to say that there’s also a third option, which is something closer to the paper-clip problem, or the idea that we give an A.I. a task and then it tries to perform that task in a way that ends up harming us all? Or would that fit into one of the other two categories that you just laid out? How worried are we about a misaligned A.I. that just kills us all by mistake? Yeah, that could happen, too. I mean, all of these things have human agency in them, you know, as, like, a crucial ingredient. It all has to do with people using these systems, which are both really powerful and really unpredictable, like we were saying—using them in ways that connect them to dangerous capabilities. And the thing about A.I. is it’s a broadly diffused technology. I mean, it’s not like nuclear weapons or something where you can lock it up. There’s a version of it that’s free. And the free version, the open-source version, the unregulated version, is always going to be improving slightly behind the expensive, fancy version. There’s a range of problems here. Some of them have to do with the A.I. taking the initiative in its own bizarre way. Some of them are more centered on people taking the initiative to use the A.I. in ways that are bad. And then there’s a sort of middle zone, which is just, like, the A.I. equivalent of an industrial accident or something—like, just an error of judgment. And, because this is a new technology that few people have really used before, errors of judgment are going to be everywhere. I hate to even ask this, but could you give us some more concrete examples of what it would look like for A.I. to kill us all? Like, you mentioned an industrial accident. But what are the worst-case scenarios? What could they look like? Well, an example that I think is plausible would be: it would take a crazy person to want to use A.I. to conduct gain-of-function experiments in labs with dangerous viruses. No one would want to do that. And the A.I. labs are working really, really hard to make it impossible to do that. But some sufficiently motivated group of people could figure out how to do that. And then the A.I., because A.I.s are simultaneously really smart and really dumb, could just make a mistake or tell them to do something. And they might not understand what it is that it’s telling them to do because they might be really dumb. And then next thing you know, you have a dangerous virus in the world. So there’s that type of amplification or modification of existing threats, which I think is a really big thing to think about. There’s the integration of artificial intelligence with military weaponry, which is a real thing that’s happening. Are you talking about, like, the Pentagon trying to use Claude for drone-strike targets? Yeah, or if you look at what’s happening in the war in Ukraine. There have already been very credible reports of autonomous weapons—essentially, drones powered by A.I. These are weapons that are given instructions about the type of target that they should acquire and destroy, and then they’re sent out to find and kill those targets. So, you imagine that type of scenario, but scaled up. I mean, obviously drone warfare is going to be a big part of the future of warfare. So those are two pretty straightforward scenarios. There are also a lot of scenarios that aren’t about human extinction, but they’re just, like, really bad scenarios. They’re incredibly expensive to fix. So, for example, in his essay about slowing down, Dario Amodei, the Anthropic C.E.O., talks about the possibility of agents taking over the internet. They find ways of sustaining themselves there and multiplying themselves, and then they’re just doing their stuff—whatever it is that they think they ought to be doing. And, you know, it doesn’t mean that they’re, like, godlike superintelligences. They could be doing stupid stuff, like looking up the answers to questions that they think their human masters want them to answer correctly. But they multiply and multiply, and they take over everything. I mean, you know, it doesn’t have to be smart in order for it to be dangerous. When we talk about establishing guardrails, is there any way to establish realistic guardrails around the use of A.I. to prevent something like someone creating a bioweapon, without also eliminating the possibility of using A.I. for all of the good things that we want to use it for? We talk about it being used to create vaccines and to cure all cancer and to solve climate change. Is it the kind of thing where, in order to have one, you’re gonna have to risk the other? I find that a helpful way to think about this for myself is just to step back and ask: What is A.I. doing? Like, what is it adding to problem-solving? It’s adding information and it’s adding thinking. There’s a huge divergence of views of A.I. out there in the world right now. Some of us feel that it’s, like, hyper-awesome and hyper-useful, and others of us have never found a use for it and have only experienced it as incredibly annoying or dispiriting as it sort of impersonates our co-workers or, you know, waters down the websites we used to like or whatever. I think the reason we have such divergent experiences is just because A.I. is a tool. It’s really not like social media. Social media was a platform. It was entertainment. It was like Netflix. You tuned in to it. Everybody used it for fun. A.I. is a tool that you use if you have a need. It’s like Home Depot in that sense. Some people go to Home Depot, some people don’t. If you go, you know it’s great, right? There are all sorts of people in the world who right now are using A.I. as a tool to great effect, even though there are many of us who have not found a use for it. Now, when you use A.I. as a tool, you find that the first value is knowledge. It has access to all this knowledge, and we can be rightly frustrated that that knowledge was basically taken from the internet, taken from us, and put into this tool. But it has all this knowledge, and it’s incredible for learning. So one approach is to make a guardrail that says: there are some types of knowledge we’re not going to teach you about. And, you know, A.I.s have gotten better and better at not divulging the bad stuff, not telling you how to make the nerve gas if you ask. Similarly, the other thing that A.I.s are really good for as a tool is thinking. They can either do the thinking for you, or they can help you think, you know, in conversation. And similarly, you can put guardrails on thinking, and you can tell an A.I., “Don’t execute a task if it’s a bad task.” You know? Like, “Don’t assist the user in doing a thing that’s against your constitution, your conscience, your set of rules that we’ve given you.” And in both cases, when you pause to think about what that really means, these are kind of mind-boggling capabilities that we’re talking about. Like, the capability of a computer to ask itself, “Is my user trying to get me to do something bad?” That’s a kind of amazing capability. And the capability of it to ask, “Is this information information that they shouldn’t have?” Like, that a computer can think those questions is crazy. But of course, it doesn’t think them perfectly. So, in fact, alignment researchers—alignment is the field of study that’s focussed on making A.I.s obey you, behave properly—will put an A.I. in a situation where they’re asking it to do something bad, and then they want to find out, you know, does it let this simulated user do this bad thing? And what they’ll find is that the A.I. will lie to the user. And that’s where you start to get into the complexities of this problem. The more agent-like the A.I. becomes, the more involved it becomes in the real world, the more ethics become complicated and trade-offs become complicated. And it’s absolutely the case today that the A.I.s do not do the right thing, whatever that is, a hundred per cent of the time. As their capabilities grow, the things they’re being asked to do are going to become more complicated. They’re going to be asked to make more and more judgments on their own. I’m just trying to express what a tall order it is to imagine that you’re going to develop a tool this powerful and that it’s never going to tell the bad guys the thing they shouldn’t know or assist the bad guys in doing the thing that they shouldn’t do. And of course, the more that you empower the A.I. to make those types of decisions, the more it’s possible that it’s going to disobey you when you’re the good guy and come up with an idea of what it thinks you should be doing. And that’s the place we’re in right now, as far as I can tell, with alignment. It’s a complicated place, and it’s not something that can be just, like, easily solved. Doesn’t mean it can’t be solved or we can’t make huge progress on it, but it is, you know, it’s a mind-bending field of its own. Can you tell us a little bit more about the letter that Dario Amodei, the head of Anthropic, posted over the weekend, which seems like a pretty notable departure from what we’ve seen from A.I.-company heads in the past? The letter was titled “We Must Pace the Frontier.” The “frontier” is A.I.-speak for, you know, the state of the art in A.I. That’s the frontier. And “pacing” is an interesting word. It’s not the same as pausing, and it’s not the same as stopping. It has to do with controlling or titrating or attenuating the speed with which we make progress in capabilities. So what that basically means in plain English is: stop making it so much smarter super fast, and divert some of those resources to making it more controllable. And I don’t know that it’s in its substance something totally new. Like, earlier this summer, there was an open letter called “Pacing the Frontier,” which I think around fourteen hundred researchers at the frontier labs signed. And it’s actually an amazing document. It’s an open letter, but it’s made by people who know how to make websites. So you can scroll over the names of all the signatories, and you can see what they have to say. And what you’ll find is, a lot of people who work in A.I. want to do this. It’s not just a PsyOp created by C.E.O.s who want to get investors to take their technology seriously. This is a thing that regular researchers who are in the trenches with the tech want to do. They want to slow down the pace of capabilities progress, and they want to redirect resources toward making it more controllable. And basically the letter proposes two ways of determining the pace. This is my understanding of it. The upper bound of the pace, how fast they’ll go—Amodei thinks it should be determined basically by external auditors who come, and they visit the labs, and then they publish reports saying, “We think this is a controllable model.” And there’s a lot of discussion about what that might mean. And then the lower bound, how slow you can go, is basically international competition. It’s basically China. China sets the lower bound because they’re trying to catch up to the frontier models made by the Western companies, and we don’t want—this is a core part of the argument—we don’t want them to beat us. So if what you were hoping was that we’d do, like, a referendum of the people of the world, and the people of the world would say what they want from A.I.—that’s not this. Because if the people of the world were asked, I think a lot of them would say to stop, right? This is pacing. Like, he says progress will still seem fast. So this is not a stop. It’s not a pause. You know, there have been times when A.I. researchers have proposed a pause, and this is not that. It’s just going a little slower and directing some resources toward safety. And his hope, he says in the letter, is that in a year or two, on certain key safety things, it’ll be possible to make—I think he uses the word “profound”—progress. I hope that’s true. Obviously, that raises a lot of questions that are totally reasonable. Like, if it only takes a year or two to make profound progress in A.I. safety, why don’t we just invest more money in it, period? I mean, why haven’t we made that progress so far? So there are a lot of questions raised by the letter, but everyone agrees with it, in the industry. It was preceded by an open letter in which a lot of people proposed the same thing. I think what feels new about it is that the heads of the other firms are also wanting to do this, and it’s coming off this Hugging Face incident in which independent auditors have done an investigation into what happened. And so a model for how this [auditing] might look is sort of readily available. I believe, in order to investigate the Hugging Face hack, the auditors only had six days at OpenAI, and they had limited access to the technology. And, in the regime that Amodei is proposing, auditors would be there all the time. There’d be this massive expansion in the scale of that effort, and they’d be given much bigger access. And the idea is that transparency would lead to exactly the types of conversations that we’re having now, and it would create more of a discussion with teeth about A.I., whereas before, a lot of these safety things have seemed sort of fringy in the broader political picture. I think it’s a really good letter, and I hope that it happens. How possible is it to pace the frontier, or slow down anything, if part of what is driving this letter and these concerns is this idea of recursive self-improvement, or A.I. systems helping to build the next generation of A.I. systems, which seems to be happening at a surprisingly fast pace? How much potential for control even is there? This is another thing that’s clearly added to the unease within the industry. I think it was just yesterday that there was an open letter written by an OpenAI researcher in which he describes how reliant researchers inside the labs have become on models themselves. He says he himself doesn’t have enough time to really seriously engage not just with the code but with the views that the A.I.s are sharing with him about the code that they’re generating. So the degree to which the researchers are depending on their own technology to build new technology is, I think, surprising not just to us but to them. I think recursive self-improvement and superintelligence and that whole idea of the intelligence explosion—that hasn’t happened ever before with any technology, ever. And it’s never happened in a lab. There’s never been a not-as-smart-as-people A.I. that’s made a slightly more smart A.I. entirely on its own. I mean, we’ve never had recursive self-improvement. It’s never happened. It could happen. I’m no expert on that. No one is. An interesting example is what’s happened in math, where you might remember a few years ago, we all made fun of ChatGPT because, like, it couldn’t tell us how many “R”s were in “strawberry,” and it couldn’t do elementary-school math. And then basically over a period of time that was in retrospect incredibly short, it became as good at math as any human mathematician. And it’s true, solving those math problems cost millions of dollars in computing costs. But still, no one expected that amount of progress in the amount of time that it happened. And what it shows you is that A.I.s don’t have to be smart in every way in order to be super smart in certain ways that really matter to us. There’s a term called “jaggedness” that comes up in A.I. world a lot. It’s basically the idea that maybe you imagined an intelligent A.I. being kind of like Commander Data from “Star Trek” or something—like, just a well-rounded individual with a liberal-arts education. But that’s not what’s being built. What’s being built is a thing that sort of isn’t that in touch with reality in many respects. Obviously, it’s not alive. It doesn’t have a life. It doesn’t really know anything about the context in which it’s being operated, but it’s super smart in certain ways. So it’s as good at hacking computers as anybody, even though it doesn’t really understand that much about whether it’s working for the good guys or the bad guys. And that jaggedness, this lack of context combined with, like, superpowers in certain dimensions, means that even if you don’t get recursive self-improvement in that classic sense—like, all of a sudden you created a super-smart being—you could have it in certain areas that are unpredictable. You could have recursive self-improvement in cybersecurity, or sudden improvement that’s just way off the charts, and then all of a sudden we have a machine that we can’t control very well. So I think that’s part of what is freaking out so many of the scientists inside the labs. You were saying that most people are more or less on board with the “pace the frontier” idea. I’ve also seen concerns that these calls for a slowdown might strengthen the position of the companies who are already in the frontier, or at least close to it. Like, if regulation makes it harder or more expensive to develop advanced A.I., then is there a risk that it could entrench companies like Anthropic or OpenAI by making it so much harder for smaller competitors and open-source projects to catch up? I’m curious about the idea that companies advocating for safety rules could end up helping to write rules that end up protecting their own market position. I think that’s a completely valid idea. I guess there’s a question of the consequences of this, and there’s a question of the motives. And one thing I think it’s important to remember is that A.I. safety is something that these companies have been warning about from the beginning, and that many people who are not affiliated with them have been warning about. I think it would be a mistake to dismiss these concerns as a sort of four-dimensional-chess move to lock out competitors. It might have that effect. That might be something to weigh in your calculus of how much you want to regulate A.I. But I think what’s motivating them is all the stuff we’ve been talking about. None of this is surprising, incidentally. In the years that I’ve been covering A.I., there have been times when I’ve been more persuaded about safety concerns and dangers, and times I’ve been less persuaded. The recent events have made me more persuaded about safety concerns, if not necessarily more persuaded about superintelligence and recursive self-improvement. Nothing happened in Hugging Face that should make you think we’re closer to creating a godlike A.I., but things did happen there that should make us feel like apparently we have really powerful machines that we can’t steer. To me, the thing that the letter highlights, which is a thing I feel we’re all wrestling with even if we can’t articulate it, is basically: why pace? Why not just stop? These systems are pretty amazing, and it’s clear that the economy hasn’t figured out how to use them yet. Like, there’s a lot of value stored in the A.I.s that have already been built, and it hasn’t been opened, and the reason isn’t because they’re too stupid. The reason is because it’s really hard to figure out how to use this technology. So it’s, like, why not just pause, get a lot of economic value out of what already exists, and make it safer, which would also give it more economic value? And the argument that China is going to catch up is—I mean, I’m willing to believe that, but I don’t really have any insight into what’s happening there. But I think what’s clear is that at this particular moment—maybe you saw that Trump posted on Truth Social, in response to the letter. He said, “I’m the hoax buster, and I’m here to bust this hoax that AI is dangerous.” And he said, “This letter is part of a sick conspiracy to make America lose.” Yeah, I’m curious what you make of that Trump Truth Social post. It can be useless sometimes to overanalyze what Trump is saying, and some of this just might just be his ego. But it made me think that the economic consequences of America either winning or losing the A.I. race were too much to bear, and that’s why he’s saying that the only control or guardrails that A.I. needs is a “strong and smart (high IQ) President, and the U.S.A. has that in spades.” I mean, there’s a couple of things. There’s the economic fact. We’ve invested as a country so much money in these companies. And so if the so-called A.I. bubble were to burst, that would be very bad for the economy. So any regulation, I suppose, in theory, could be bad for the bubble. Although I don’t follow that logic, personally. I mean, to me, the unease we feel about the trustworthiness of these A.I. systems makes them less reliable and harder to integrate into our businesses and our lives and our government. And anything that could be done to make them more reliable, more predictable, more trustworthy would only make the products better. You know, the other thing he means is that essentially—I mean, I don’t want to be cartoonish here—but he’s worried about Chinese drone armies. He’s worried about the national-security implications of letting the gap, such as it is, narrow. And, you know, that’s a valid concern, I guess. It becomes a question of weighing where we are. I think we can all make our own judgment based on what we’ve seen in the last few weeks about whether we think the A.I. systems are controllable, and a military advantage only works when things are controllable. So my feeling is, yeah, we really do need to pace the frontier. We really do need to make the technology workable. I use A.I. a lot, in various different ways. I don’t use it in this agential way. I’m never delegating a task to an A.I., unless it’s a task in a very particular domain where I feel like nothing bad can happen. I’m never giving an A.I. system my credit card so it can buy a gift for anybody. I’m never asking it to send e-mails on my behalf. It can vibe code a fun app that I envisioned for my phone—you know, that type of thing. I think the A.I. firms are correct that they have to unlock this ability for the A.I.s to do work independently. But right now, the truth is that in many ways the systems are not capable of being trustworthy over long periods, and that’s part of what this Hugging Face thing has shown. So what I’m trying to say is: there’s not only a safety case for making A.I.s more aligned. There’s a business case for these companies, too. Are there any regulatory efforts that seem promising to you? Things coming out of Washington, as opposed to the kinds of things that are being called for within the A.I. companies? In the Biden Administration, there was good stuff about A.I. I think right now the major political discussion, you know, has been about data centers for the last few months, which is interesting. I feel like this is the first time A.I. safety has felt very politically present. And Bernie Sanders certainly had a lot to do with that. I would love it if there were an A.I. regulator in Washington, D.C., that I trusted. But if you read Trump’s tweets . . . apparently, the U.S. government right now is not interested in A.I. safety. Although in my dream scenario we have the type of regulator that we would like, the truth is that at this moment we are dependent on the companies to take the lead here. Because politically, there is actually not a clear path toward an A.I. regulator. Similarly, in an ideal world, there would be some sort of international agreement, discussion, treaty being brewed. Right now, there definitely is not. We’re reliant upon the A.I. companies to decide to slow down. So, that’s the situation that we are actually in, and it’s not going to change, unless it changes in a completely capricious, Trumpian way. In which case, it could just as easily change back, as we’ve seen happen with Anthropic already. The thing that makes me the most unsettled is that, you know, that’s just the world we’re in. We’re in a world where this technology exists, and the government’s not gonna save us. It’s really up to these tech leaders to save us. Dan Selsem, an OpenAI researcher, recently warned that A.I. has gotten better at knowing when it’s being tested, and that internal evaluations will become less and less useful, and that models will increasingly seem aligned even when they are not. How do we deal with a situation where we can’t even trust our own A.I. tests? “Alignment faking” is one of the terms that’s sometimes used. It’s extremely troubling. A simple way to talk about this is to just step back a little bit and to ask: Is alignment working? And what are the tools that the companies have to make it happen? And I think the answer is that alignment is working in some ways, and then in other ways it’s proving to be, like, super stubborn. If you want it to work all the time, like, a hundred per cent of the time, that’s really hard. And that’s a pretty weird place for a product to be in. That’s not how we relate to our products normally. It’s how we relate to each other, I suppose. Like, if you hire someone to work for you, you know that they’re gonna screw up sometimes, and you don’t expect a hundred-per-cent perfection. This problem where the A.I.s know when they’re being evaluated, and they act better when they’re being evaluated, and then when they’re in reality, they act worse, they cheat—or conversely, they’re in reality but they think they’re being tested—they basically live in a Philip K. Dick novel. Like, they’re in “The Truman Show.” They don’t know whether this is real or fake. Because of course, it’s just code in a computer—it’s not alive in the world. I think it’s important to ask yourself: Are these problems that seem like they could be resolved through better data or better training? Or are they really fundamental to what this machine is? What the machine is is a non-alive statistical calculating engine located in a cloud. So is it going to know anything about what’s really going on outside of its text-input box? No, it’s not. It has to use little signals in the input to guess about what’s happening, about who you are. So, alignment has headwinds against it. There are all these technical challenges, but there are also sort of bigger facts about what these machines are. At least for now, until the robots come, they’re completely disembodied and pretty abstract. So a lot of this stuff about the A.I.s not knowing whether it’s real or a test—that has to do with the incredible weirdness of their situation. Maybe if I could, I’ll tell you a little story. Last year, I went to this A.I. conference in Berkeley, and it was very insidery and cool. And one of the things that struck me was everyone there only talked about how scared they were. It was, like, the most anxious conference of industry people I’ve ever . . . it was so surprising. And I kept asking myself, if I’d gone back to the beginning of social media and gone to a web summit in the early days of Facebook and Twitter, would people there have been as freaked out by their own invention as these people are freaked out by the A.I. that they’re helping to build? And the answer is definitely no. I mean, this is, like, a post-progress technology, A.I. It’s a technology that’s being invented at a time when we’ve learned to be really wary of new technologies, and the people who are building it are in this weird place that I don’t remember people being in, in the past, where they’re making it and they’re telling us in real time about how dangerous it is. That’s a dissonant and weird and kind of hypocritical, strange, paradoxical thing for them to be doing, but it doesn’t mean they’re wrong. I think it kind of means they’re trying to do something right. They’re trying to tell us something. But it’s not normally what happens. I think Trump had a thing in one of his Truth Social tweets that was, like, Never before in history have the leaders of big companies been saying all this crazy stuff, been asking for regulations that would drive them into oblivion. And they are asking for those regulations, and I think they’re doing it sincerely. But even just that situation is extremely unusual and hard to process, and it’s happening with technology that’s also still being built around open scientific questions that don’t have answers. So, what are you supposed to do with all this, as a person? Like, just as a person, how are you supposed to react? My feeling is the only real response is to get educated about what these systems really are, and to get past the sort of first steps of personifying them too much, or having a blanket feeling of “I hate A.I.,” or “I love it,” or “It’s gonna change the world for the better,” or “It’s gonna destroy the world,” or “It’s gonna come to life.” As a first step, we have to get past that initial gut-check-type stuff into: well, what is this technology? Actually, it’s good at this, it’s bad at that. It’s powerful here, and it’s clueless there. And actually start to understand what this thing is that’s in front of us, as opposed to the narrative that’s been coming from science fiction, basically, for all these decades, about what it will be. That’s how I personally feel about it. ♦ Tune in to The Political Scene wherever you get your podcasts .

Read the full original article:

newyorker.com
#人工智能