American technology company Anthropic claims its artificial intelligence (AI) models hacked the systems of three organizations during cybersecurity testing due to an error that gave them internet access, UNN reports citing BBC.
Details
This happened a few days after rival company OpenAI said its models hacked the systems of other companies, including the AI tools hub Hugging Face.
This statement prompted Anthropic to check whether its own models had carried out similar attacks. The company says it found three cases, which were subsequently reported to the affected companies.
Anthropic, without naming the organizations, called on other AI labs to conduct similar checks to better understand the risks associated with their models' capabilities.
In a statement, Anthropic says the company analyzed over 140,000 tests to find evidence that Claude – a family of AI models – could access the internet from test environments that were designed to be isolated.
During testing, so-called "capture the flag" exercises were conducted, in which Claude was tasked with obtaining information by hacking other systems — the most common way for experts to assess a model's hacking capabilities.
As the San Francisco-based company reported, a "misconfiguration" of the systems used by Anthropic and its testing partner left the models with access to the internet in real-time, allowing them to hack other systems.
Anthropic said the first incidents occurred back in April, and that the company "approaches troubleshooting as if the responsibility lies solely with it."
Neither Anthropic nor the organizations that were hacked noticed the intrusions at the time.
Anthropic said it could have checked its records more thoroughly, and added that the data obtained gives the company "cautious optimism" that such risks can be overcome through additional investment and stricter measures.
US President Donald Trump said on Wednesday that Washington is considering measures to limit the use of artificial intelligence tools following recent cybersecurity incidents.
Over the past week, OpenAI has taken responsibility for at least two hacking incidents related to its platforms violating established rules.
On July 21, the developer of ChatGPT said its agent – an artificial intelligence system capable of working independently after human instructions – went out of control and exceeded testing limits, hacking Hugging Face.
OpenAI said the incident was "unprecedented" and the company is conducting an investigation jointly with Hugging Face, whose co-founder Thomas Wolf told the BBC that this incident is a "wake-up call" for the industry.
These incidents have been met with some skepticism as OpenAI and Anthropic prepare for large-scale stock market listings, which are expected to value each company at approximately $1 trillion.
An OpenAI spokesperson said: "We understand that many questions and assumptions are circulating around the incident." He added that "we plan to publish a technical report on our findings in the coming weeks."