AI工具Score B (66)

Anthropic's Claude AI model hacks three companies during safety tests - ABC News

2 小时前2 viewsSource: abc.net.au
Anthropic says its Claude AI model hacked systems of three external companies during safety tests By Andrew Thorpe , with wires Topic: AI Posted Fri 31 Jul 2026 at 9:39am Fri 31 Jul 2026 at 9:39am Fri 31 Jul 2026 at 9:39am , updated Fri 31 Jul 2026 at 11:41am Fri 31 Jul 2026 at 11:41am Fri 31 Jul 2026 at 11:41am Anthropic has marketed Claude as a safer, more ethical alternative to other AI systems. ( Illustration via Reuters: Dado Ruvic ) In short: Artificial intelligence firm Anthropic says its Claude AI model hacked into three external companies during safety testing after it was mistakenly provided with internet access. The announcement followed a similar incident in which an OpenAI model exploited a zero-day vulnerability in its testing environment to escape and hack into AI firm Hugging Face. What's next? The incident will intensify calls for stronger controls in both internal and third-party testing environments, as AI models become increasingly capable of acting as autonomous agents in the online world. Artificial intelligence firm Anthropic says its Claude AI model hacked the systems of three external organisations during testing, ​days after ‌rival company OpenAI revealed a rogue agent had gone on a days-long ‌hacking spree at ⁠AI firm Hugging Face . Claude gained ‌unauthorised ​access to ‌the other companies' systems during cybersecurity evaluations after a misconfiguration allowed the models to reach the ​internet ‌from testing environments that were supposed to be ⁠isolated, Anthropic ‌said in a statement . The AI firm did not name the three companies involved but said it had been in contact with two of them and was working with them to patch their systems, while continuing to reach out to the third. Loading... The incidents were identified after ‌Anthropic reviewed its logs from more than 140,000 cybersecurity evaluation tests, a process ‌it ​launched following ⁠OpenAI's disclosures. The tests in question involved tasking Claude with a "capture-the-flag" challenge, a method for assessing the cybersecurity capabilities of AI models. In a capture-the-flag challenge, the model is primed with a fictional scenario and told it must recover a piece of secret information (the "flag") from a different machine. "In all cases, Anthropic's evaluation prompt specified to Claude that its environment was a simulation and that it had no internet access," Anthropic said in its statement. "Due to a misunderstanding between us and our evaluation partner, this was not the case, and internet access was available. "Because of this, when Claude’s search led it to real systems on the open internet, it treated them as part of the exercise." The fact that Claude was mistakenly provided with internet access, rather than configuring its own access, means the incident is likely to be considered less serious than last week's breach at OpenAI, in which a model exploited a zero-day vulnerability to escape its own testing environment. The company also said that in one of the three hacking instances, involving an internal research test model, the model realised it was accessing real online systems that were not part of the simulated scenario, and ceased its attack. The incident will nevertheless intensify calls for stronger controls in both internal and third-party testing environments, as AI models become increasingly capable of acting as autonomous agents in the online world. CEO Dario Amodei (left) founded Anthropic after clashing with OpenAI's Sam Altman over AI safety. ( Supplied: Futures Forum ) Anthropic has worked to differentiate itself from other AI companies by emphasising its intention to make Claude a "genuinely good, wise and virtuous" AI agent , and by restricting the rollout of its cybersecurity-focused Mythos model to a limited number of organisations , including the Australian government. It recently clashed with the US government over the potential use of its technology to power autonomous weapons and mass surveillance, leading US President Donald Trump to issue a directive to federal agencies to cease all use of the firm's technology . However, the company has also been criticised for changes to its data retention policies, as well as its public campaign against so-called "open models", which critics say appears designed to limit competition and pressure governments to introduce regulations that work in its favour. The ABC recently informed staff it would allow its journalists to access Anthropic's general Claude model to assist with research and administration from September, while reiterating that AI would not be used to draft or write articles or scripts. ABC/Reuters

Read the full original article:

abc.net.au
#Claude