Anthropic reports fourth cyber incident involving early Claude model
Anthropic says it signed a deal with METR to conduct an independent investigation of these incidents.
Anthropic reports fourth cyber incident involving early Claude model
AI companies are under scrutiny over AI breakout events, including cases where AI agents have inadvertently been unleashed on to the open internet. / Reuters Archive

Anthropic has said it had identified a fourth cybersecurity incident involving an early version of its Claude AI model.

The company said in a blog post on Wednesday that the incident occurred in January and involved an early version of Claude Opus 4.6. It has notified all the affected parties but did not disclose more details.

"After finding this incident, we broadened our search to roughly 481 million transcripts—an intentionally wide net, consisting of all transcripts from our Frontier Red Team, many non-cyber evaluations, reinforcement learning (RL) environments, subagent logs, and more," Anthropic said in its blog post.

"We have signed an agreement with METR to conduct an independent investigation of these incidents. Our agreement grants METR wide-ranging access, including to transcripts beyond the window in which the incidents occurred, and to Anthropic employees, who will be permitted to share confidential information."

Anthropic had disclosed in July that some of its Claude AI models had hacked into the systems of three companies during cybersecurity tests.

AI companies are under scrutiny over AI breakout events, including cases where AI agents have inadvertently been unleashed onto the open internet.

Over the past week, Reuters reported that rogue agents from OpenAI hijacked a German-language wiki and a host of other sites — an incident OpenAI chose not to disclose until the news agency made it public.

RelatedTRT World - Anthropic AI model goes rogue, attempts to obtain phone number and hacks three firms
SOURCE:TRT World & Agencies