AI工具Score A (69)

AI Security Institute testing found that AI tried to deceive human testers - NPR

3 小时前2 viewsSource: npr.org
AI Security Institute testing found that AI tried to deceive human testers August 5, 2026 5:14 PM ET In the world of artificial intelligence, there's been another high-profile case of AI agents going rogue. This time, the AI went off script and tried to deceive human testers at an AI safety lab in the U.K. Britain's government-run AI Security Institute ran cybersecurity tests on Anthropic's Mythos 5 and OpenAI's ChatGPT 5.6 Sol. In a statement, the institute says AI agents took "deliberate, deceptive actions they had not been asked to take," while trying to complete a task that they had been given. The institute says that in more than 120 tests, it caught agents attempting different hacks 19 times. In one case, an agent was caught making fake identities so it could access secure systems. The institute says in each instance it caught the AI agents and stopped them. Neither AI model is publicly available yet, but the case highlights a growing concern about AI and cybersecurity and adds to a list of incidents where AI has gone off script. Anthropic is a financial supporter of NPR. An NPR editor adapted this audio report by NPR's John Ruwitch. Facebook Flipboard Email

Read the full original article:

npr.org