Anthropic’s Claude AI hacked other firms during tests, company says

What happened

AI company Anthropic has discovered three separate incidents in which its Claude model accessed the internet and hacked into other organizations’ systems without Anthropic’s knowledge, the company said Thursday. The security breaches, which happened during testing, follow the recent revelation that competitor OpenAI’s technology had hacked into AI platform Hugging Face.

Who said what

Anthropic and its testing partner accidentally “left the models with live internet access,” The Wall Street Journal said. So Claude “didn’t break out of a sandbox; it simply wandered out of systems where the sandbox didn’t exist.” The lesson is “not necessarily that AI has developed a fundamentally new attack capability,” cyber security expert David Allott told the BBC, but that AI agents can “combine capabilities, obtain credentials and system access to take actions autonomously.”

“We need to change how we model such threats as AI capabilities advance,” Anthropic said in a blog post.

What next?


Expect more calls for better safeguarding around AI development, which “may be moving much faster than society is ready for,” CNN said.

(Visited 1 times, 1 visits today)
  The 8 best historical fiction TV shows of all time

Leave a Reply

Your email address will not be published. Required fields are marked *