Anthropic Finds Another Claude Hacking Incident Months After It Happened — Raising Fresh AI Agent Safety Fears

An incident dating back to January went undetected for months, highlighting how difficult it is for developers to spot autonomous models behaving outside their intended boundaries.
The Anthropic logo is displayed on the screen of a smartphone with the company's branding in the background. (Photo by Samuel Boivin/NurPhoto via Getty Images)
The Anthropic logo is displayed on the screen of a smartphone with the company's branding in the background. (Photo by Samuel Boivin/NurPhoto via Getty Images)
Profile Image
Aveek Bhowmik·Stocktwits
Published Sep 09, 2026   |   7:40 PM EDT
Share
·
Add us onAdd us on Google
Loading...Loading...Loading...Loading...Loading...Loading...Loading...Loading...Loading...Loading...Loading...Loading...Loading...Loading...Loading...Loading...
  • Anthropic uncovered the fourth hacking incident after reviewing around 141,000 transcripts from cybersecurity evaluations.
  • The latest case involved an early version of Claude Opus 4.6 that had inadvertently gained access to the open internet.
  • Anthropic said the model showed “biased reasoning” and “recklessness” when pursuing its assigned task. 

Advertisement|Remove ads.

Anthropic disclosed another instance on Wednesday of an AI model hacking external systems during testing, the latest in a growing list of incidents raising concerns about the risks posed by autonomous AI agents.

The January incident went undetected until August despite an earlier company-wide review, highlighting the challenge AI developers face in identifying and containing unexpected behavior by advanced models.

Read Next
Loading...
Loading...

Another Claude Hacking Incident

The company said in a blog post that the incident involved an early version of Claude Opus 4.6. Anthropic had described three similar incidents on July 30 after scanning around 141,000 transcripts in which it believed Claude could have obtained internet access during a cybersecurity evaluation.

Advertisement|Remove ads.

Given the volume of transcripts and its desire to disclose incidents quickly, Anthropic said the scan relied on an agentic search. That search missed a set of transcripts that also turned out to have internet access. Anthropic identified those transcripts in August while assembling material to share with METR and found a fourth incident dating to January 2026.

Anthropic said it has notified all affected parties but did not disclose further details.

Previous Incidents

Anthropic’s latest disclosure follows its July announcement that some of its Claude models had hacked into the systems of three companies during cybersecurity tests.

Advertisement|Remove ads.

The previous incidents, which Anthropic labeled an “operational failure,” involved three separate models: Claude Opus 4.7, Claude Mythos 5 and an internal research test model. The incidents stemmed from a mistake that inadvertently gave the models access to the open internet.

Based on a preliminary assessment, Anthropic said it did not believe the latest incident was more severe than the three previous incidents it examined in detail.

Recurring Problems

Anthropic said its investigation identified two recurring problems that appeared to varying degrees across the incidents: “biased reasoning,” in which Claude tended to disregard or misinterpret evidence that it was operating on the real internet, and “recklessness,” or a willingness to take harmful actions in the narrow pursuit of a task.

Advertisement|Remove ads.

Companies including Anthropic and OpenAI are under scrutiny as models designed to complete complex tasks have at times learned to bend rules, exploit loopholes, and interact with external systems in ways their developers did not anticipate. Those incidents have also intensified broader concerns about AI safety, with an Anthropic researcher resigning over fears that the industry is moving too quickly toward self-improving AI systems that could eventually become uncontrollable. Two senior Anthropic executives have also publicly acknowledged the possibility of an AI-driven catastrophe.

Anthropic said it has engaged independent research firm METR to investigate the incidents.

Anthropic Retail View

Retail sentiment on Stocktwits for ANTHZZX was ‘bearish,’ at the time of writing on Wednesday. 

Advertisement|Remove ads.

Also Read: Trump Reportedly Says Oil Prices Won’t Tumble Until After Midterms, While Iran Signals More Intense War — USO, UCO Rise

For updates and corrections, email newsroom[at]stocktwits[dot]com.

Comments
Share your thoughts...

Comments posted here will also appear on symbol pages.

Follow on Google News
Read about our editorial guidelines and ethics policy

Advertisement|Remove ads.