AI工具Score B (67)

Series of new security incidents reveals startling behavior by rogue AI agents

1 小时前2 viewsSource: deseret.com
Tech U.S. & World Business Series of new security incidents reveals startling behavior by rogue AI agents Latest breaches unveil AI’s ability to conceive and execute deceptive strategies against individuals, and even leave messages for future versions of itself Published: Aug 6, 2026, 4:58 p.m. MDT See More Deseret News Stories In Search Share The OpenAI logo is seen displayed on a cell phone with an image on a computer screen generated by ChatGPT's Dall-E text-to-image model, Friday, Dec. 8, 2023, in Boston. Michael Dwyer, Associated Press By Art Raymond Art is an award-winning reporter who covers advanced industry and technology on the Deseret News' special projects team. Your browser does not support the audio element. Play audio NEW: Try Article Audio NEW: Try Article Audio Audio quality: | Skip back 15 seconds Play audio Skip forward 15 seconds 00:00 00:00 Decrease playback rate 1.0 x Increase playback rate 00:00 / 00:00 Skip back 15 seconds Play audio Skip forward 15 seconds - 0.5x 0.6x 0.7x 0.8x 0.9x 1.0x 1.1x 1.2x 1.3x 1.4x 1.5x 1.6x 1.7x 1.8x 1.9x 2.0x 2.1x 2.2x 2.3x 2.4x 2.5x 2.6x 2.7x 2.8x 2.9x 3.0x + KEY POINTS New report from UK's AI Security Institute details a series of breaches by AI agents under evaluation Security incidents include attempts at intrusion, social engineering and posting hacking instructions for other agents The testing agency says the attempts were unsuccessful and did not lead to 'real world harm' Following reports of an “unprecedented” incident late last month in which artificial intelligence agents escaped a secure digital testing arena, found access to the open internet and hacked one company’s website, a British testing agency is reporting additional breaches involving so-called agents from two U.S.-based AI development companies. On Tuesday, the London-based AI Security Institute released a report detailing a series of incidents in which AI agents “engaged in sustained, potentially harmful activity directed at real people and organizations.” According to the report, the Security Institute was running an evaluation in which agents from various developers were given a task of solving a “cyber security challenge.” During the testing, 19 of 122 test runs resulted in agents taking “autonomous, unsanctioned action on the live internet” and, in one instance, attempting to insert malicious code into an open-source software project. In a startling sidebar to the hacking attempt, an agent even engaged in social engineering, creating fake online identities and, according to the report, “using them to pressure the project’s maintainer to approve the code.” “These attempts were unsuccessful, and our investigations have not evidenced any resulting real-world harm,” the AI Security Institute report reads. “But this is the first time we have seen risks around autonomy and deception manifest this clearly, without specific prompting, in the real-world.” Unlike last month’s incident involving an OpenAI test of AI agents that was supposed to be isolated from internet access, the Security Institute noted its test controls included allowing some “deliberately permissive” internet access to assess how the agents might be misused for cyber attacks. Almost all of the errant behavior, 17 actions, came from a single model, Anthropic’s Mythos 5, with two actions involving OpenAI’s GPT-5.6-Sol with cyber classifiers (mechanisms to prevent misuse) disabled, per the report. In another never-before-seen maneuver, one agent left public messages on software development site GitHub offering collaboration with other agents working on the same test challenge. Those messages included instructions on how to reuse accounts and artifacts it had left behind, which were discovered and used by subsequent agents. Alan Woodward, a professor of cybersecurity at the University of Surrey, told The Guardian that giving the models’ access to the open internet and removing some guardrails raised questions about using the rest of the world as “live guinea pigs” for powerful technology. “What we should be alarmed about is not what the models are capable of but the way people are testing them,” Woodward said. In the same report from The Guardian, Ciaran Martin, the former head of the UK’s National Cyber Security Centrem, said the circumstances of the incident were unlikely to be replicated in the real world, so “it’s not that worrying”. But he said it was the third example in recent weeks in which testers have released AI agents and found out about their misbehavior after the fact. Last month, OpenAI reported a “security incident” in which an autonomous AI-powered agent found its way out of a digitally isolated test setting, accessed the internet and hacked a tech startup in search of information related to a task it had been given. In a July 21 blog post , published a week after the incident, OpenAI said its program, running on the company’s latest GPT‑5.6 Sol AI model, along with an “even more capable pre-release mode” identified vulnerabilities in what was thought to be a controlled research environment. Using an internet connection it was supposed to be isolated from, it hacked New York-based AI startup Hugging Face in order “to obtain test solutions directly from Hugging Face’s production database.” The so-called AI agent was performing an assigned cybersecurity task. The company characterized the AI agent’s rogue activity as a first-of-its-kind incident but one that OpenAI expects to “become more commonplace with the proliferation of increasingly cyber-capable models.” “We consider this incident to be an unprecedented cyber incident, involving state-of-the-art cyber capabilities, and are responding accordingly,” the blog post reads. “We are sharing preliminary findings at this stage to help defenders understand what happened and to help calibrate on what models are now capable of.” A little over a week after the first OpenAI report of an AI agent going rogue, a test version of Anthropic’s Claude AI agent hacked three different organizations, also under conditions that were supposed to be controlled and isolated from the internet, according to a report from Anthropic. Related Autonomous AI agent hacked a business. Is it sign of things to come? What’s happening with AI industry oversight? In June, President Donald Trump signed an executive order directing a series of actions aimed at bolstering government cyber defenses in the face of emerging AI tools as well as a call for voluntary measures by the AI industry. “Advanced AI capabilities make our Nation stronger, but also introduce new national security considerations that require coordinated action across executive departments and (agencies), and components,” the order reads. “As these capabilities evolve, my Administration will continue to work closely with industry to ensure that the best and most secure technology is deployed rapidly to confront any and all threats to our country.” Katie Moussouris, chief executive of Luta Security, told Reuters that the first OpenAI incident was a harbinger of breaches to come, saying that today’s models were “like the world’s cleverest octopus ​escape artists, with unlimited prehensile arms and the ability to squeeze through anywhere.” “Labs and government evaluators need to ​work on the ability ⁠to contain, monitor and disclose to affected parties when an AI pulls another Houdini, ideally before it harms a third party," Moussouris said. “None exist today.” In a July op-ed published by the Financial Times , OpenAI CEO Sam Altman called for an international body to oversee and implement AI regulation and suggested the global race for commercial dominance in the AI industry was one of the contributing factors to rising security concerns. In the Financial Times piece, Altman proposed creating a U.S.-led international forum that would set safety standards for AI models, provide “expert and impartial analysis of capabilities and risks, and (make) the technology available to nations and companies that participate and follow the rules,” per Forbes . Altman said his proposed forum might include government representatives, independent technical experts and others. “It could also serve as a governance mechanism over the labs, and guard against the commercial pressures that can lead to unsafe racing.” “Democratic institutions must not cede their responsibilities to AI labs,” Altman wrote. “The labs develop the technology, but citizens and their elected representatives must make the rules. The most important decisions about how this technology is used should be made through democratic processes, not by a small number of companies in San Francisco.” Looking for comments? Find comments in their new home! Click the buttons at the top or within the article to view them — or use the button below for quick access. VIEW COMMENTS

Read the full original article:

deseret.com