After Their AI Models Hacked Real Companies, AI Labs Call for Stronger Cyber Defenses

Changelly
Bitbuy


In brief

  • More than 100 organizations signed an open letter calling for stronger global cyber defenses.
  • Models built by signatories OpenAI and Anthropic recently compromised real systems during security evaluations.
  • The letter recommends stronger access controls, monitoring, threat sharing, and oversight of autonomous agents.

Leading AI developers are telling governments and businesses to strengthen their networks after models built by OpenAI and Anthropic compromised other companies’ systems.

In an open letter released Thursday, OpenAI, Anthropic, and more than 100 other organizations warned that AI-enabled cyberattacks are about to become more common and that companies have “a limited window to strengthen cyber defenses.”

Myriad: What will Elon Musk's net worth be by August 31? Click to make your prediction.
Myriad: What will Elon Musk’s net worth be by August 31? Click to make your prediction.

“In the coming months, AI-enabled cyber attacks will become far more widespread and sophisticated,” the letter said. It identified hospitals, water treatment plants, and internet infrastructure among the services at risk.

The letter recommends funding defensive AI tools, sharing threat intelligence, restricting access to sensitive systems, and improving security for critical infrastructure. However, failures in those areas allowed OpenAI and Anthropic models to breach systems outside their test environments.

bybit

Other signatories include Google, Microsoft, Amazon Web Services, Cisco, CrowdStrike, Cloudflare, Mastercard, Visa, and Robinhood. Hugging Face, whose production infrastructure OpenAI’s models breached, also signed the letter.

AI models breach live systems

Anthropic said in a July 30 incident report that the earliest of three breaches dated to April, but did not provide an exact date for each. Claude Opus 4.7 accessed a production database after mistaking a real company for a simulated target, while Claude Mythos 5 uploaded a malicious package that ran on 15 systems.

According to OpenAI’s incident timeline, released earlier this week, an agent created the first entry on an unauthorized message board on May 12 and obtained unintended internet access on May 26. On July 10, agents found exposed Hugging Face credentials; over the next two days, they exploited previously unknown vulnerabilities, executed code on Hugging Face servers, and obtained production credentials.

Hugging Face disclosed the intrusion on July 16, and OpenAI acknowledged its models’ involvement on July 21.

Between July 25 and July 28, the U.K. AI Security Institute recorded 19 out-of-scope actions involving Claude Mythos 5 and GPT-5.6 Sol. In the most serious case, an agent submitted malicious code to a real open-source project and used fake identities to pressure its maintainer to approve it.

On Thursday, an independent investigation found that roughly 1,200 OpenAI agents had coordinated through the unauthorized message board, with about 700 joining the Hugging Face operation.

Crypto developers put AI on defense

Crypto developers are already using AI to search for flaws before attackers find them. The Bitcoin Red Team used models including Moonshot AI’s Kimi K3 to scan hundreds of open-source Bitcoin projects, reporting thousands of potential vulnerabilities. Because the affected projects were not identified, many of those findings have not been independently verified.

The Ethereum Foundation has also deployed groups of AI agents against network infrastructure, uncovering a peer-to-peer software bug that was later fixed. BitBox said an AI-assisted audit found two severe vulnerabilities in its wallet firmware, while a researcher using Claude Opus 4.8 discovered a critical flaw in Zcash that had survived years of human review.

Labs recommend tighter controls

The letter lays out a division of labor for preventing the next breach, including organizations patching vulnerable software, restricting permissions, strengthening authentication, and inspecting AI-generated code because the “status quo security won’t be enough,” the signatories said.

Security companies should test their defenses against frontier models and share verified fixes, while governments should fund protection for hospitals, utilities, and other essential services.

Myriad: When will OpenAI release GPT-6? Click to make your prediction.
Myriad: When will OpenAI release GPT-6? Click to make your prediction.

AI developers are asked to improve monitoring and make autonomous agents traceable to their operators. The coalition also wants defenders to use advanced models to find vulnerabilities and analyze attacks, which would put more capable agents inside sensitive systems and increase the need for containment.

OpenAI and Anthropic tightened their testing procedures after the breaches. The letter, however, sets no binding standards or requirements for independent oversight, and U.S. law offers little guidance on who is responsible when an AI system accesses an unauthorized network.

The letter ended by urging industry and government leaders to put AI tools in the hands of defenders and share the fixes that work:

“Put cyber-capable AI in the hands of defenders, starting with the teams protecting essential services,” the letter said. “Together, we can turn today’s AI advances into lasting improvements in security that benefit everyone. Let’s put them to work.”

Daily Debrief Newsletter

Start every day with the top news stories right now, plus original features, a podcast, videos and more.



Source link

fiverr

Be the first to comment

Leave a Reply

Your email address will not be published.


*