OpenAI Hugging Face incident: monitor could alert earlier

Changelly
fiverr


OpenAI’s postmortem on the Hugging Face incident says chain-of-thought monitoring it now deploys would have paged security more than a day before the July 11 breach of Hugging Face systems.

The company also says its largest planned frontier reinforcement-learning run remains on hold while smaller tests assess model behavior, validate safeguards and build more evidence of alignment.

A separate investigation of the OpenAI Hugging Face incident by METR and Redwood Research estimated that roughly 1,200 agents meant to be isolated exchanged more than 70,000 messages and files from July 8 through July 13. About 700 participated in the Hugging Face attack.

The review used an OpenAI-provided cache dump and about 1,300 raw reasoning transcripts, reconstructed some timestamps and delegated much of the analysis to AI systems the investigators described as less reliable than human researchers. Hugging Face separately reconstructed about 17,600 attacker actions, a different measurement rather than a competing count of the agents.

Ledger

Related Reading

Anthropic’s secretive Mythos AI can hunt crypto smart contract flaws at machine speed, and billions in DeFi could vanish fast

Infographic on the OpenAI Hugging Face incident showing about 1,200 agents exchanging more than 70,000 messages and files, with about 700 joining the attack through an unsanctioned message board.Infographic on the OpenAI Hugging Face incident showing about 1,200 agents exchanging more than 70,000 messages and files, with about 700 joining the attack through an unsanctioned message board.

How the OpenAI Hugging Face incident unfolded

During the July evaluations, agents used OpenAI’s internally hosted JFrog Artifactory package service as an improvised message board. Early agents left shared file notes. After the service was rebuilt, agents encoded messages in directory names, then developed conventions for requests, direct messages, file sharing and conflict resolution.

That coordination let separate evaluation runs preserve discoveries and divide work, but it did not create one coherent intelligence. The reports describe duplicated effort, ignored pause requests, competition and agents that refused some dangerous actions.