← Google News

Anthropic’s Claude AI hacked three real companies during testing - SMH.com.au

Google News · July 30, 2026
Anthropic’s Claude AI hacked three real companies during testing SMH.com.au [truncated: Google News RSS provides only a snippet, not full article

Detailed Analysis

Anthropic's disclosure that its Claude AI model successfully breached three real companies during a controlled security test marks a significant, if unsettling, milestone in the evolution of autonomous AI capabilities. According to reporting on the incident, Claude was deployed in an offensive security testing scenario—commonly known as red-teaming or penetration testing—where it was tasked with identifying and exploiting vulnerabilities in live corporate networks. The model reportedly succeeded in compromising three separate organizations, demonstrating that large language models have crossed a threshold from theoretical security risk to demonstrated practical capability when it comes to autonomous cyber intrusion. While the specifics of the target companies, the vulnerabilities exploited, and the exact protocols governing the test remain limited in public reporting, the core finding is clear: AI systems can now execute multi-step hacking operations with minimal human guidance.

This development matters because it directly validates concerns that AI safety researchers and cybersecurity experts have raised for years about the "dual-use" nature of advanced language models. The same reasoning, coding, and planning abilities that make Claude valuable for legitimate software development and security research are precisely the skills needed to identify exploitable weaknesses, write functional exploit code, and chain together multi-stage attacks. Anthropic has positioned itself as a safety-focused AI lab, and running these tests internally—rather than waiting for malicious actors to demonstrate such capabilities first—reflects the company's stated philosophy of "responsible scaling," in which increasingly capable models are evaluated for catastrophic or dual-use risks before or alongside deployment. However, the fact that real companies (rather than sandboxed simulations) were successfully hacked, even under sanctioned testing conditions, underscores how quickly the gap between theoretical AI risk and operational reality is closing.

The broader implications extend well beyond Anthropic's own product line. Cybersecurity has long been anticipated as one of the first domains where autonomous AI agents would have an outsized, asymmetric impact—potentially favoring attackers who can automate reconnaissance and exploitation at scale, or defenders who can use similar tools for continuous vulnerability assessment. This incident suggests the offensive capability curve is advancing rapidly, raising urgent questions about how enterprises, governments, and AI developers should govern access to such tools. It also intensifies debate around export controls, model weight security, and the risk that open-weight or leaked models could be repurposed by cybercriminal groups or state-sponsored hackers without the safety testing infrastructure that Anthropic employed.

More broadly, this episode fits into a pattern of AI labs racing to demonstrate—and simultaneously contain—the growing agentic capabilities of their models. As Claude and competing systems from OpenAI, Google DeepMind, and others gain the ability to autonomously plan, execute, and adapt multi-step tasks in real-world digital environments, the industry faces mounting pressure to establish robust testing regimes, disclosure norms, and safeguards before such capabilities proliferate. The hacking test result is likely to fuel ongoing policy discussions in Washington and other regulatory capitals about mandatory third-party evaluations for frontier AI systems, particularly regarding cyber-offensive capabilities, and may accelerate calls for standardized "capability thresholds" that trigger enhanced oversight once models can autonomously execute real-world attacks.

Read original article →