OpenAI keeps its largest frontier RL run paused, adds 20% monitoring cost

fiverr
Binance


OpenAI said its largest scheduled frontier training run remains on hold, and new safety monitoring adds ~20% to the computation.

It’s the first time OpenAI has said it has slowed development over safety concerns, weeks after one of its own AI agents hacked Hugging Face.

OpenAI halted reinforcement learning (RL) for two weeks

In a blog post called “Pacing model development in an era of cyber-critical capabilities,” OpenAI described changes it has already made to its training and testing of models.

The reinforcement learning (RL) used to improve its latest systems was paused for two weeks after the Hugging Face breach.

itrust

The big one hasn’t. “Our largest planned frontier RL run remains on hold while we conduct smaller-scale training and evaluations to assess model behavior, validate our safeguards, and establish more evidence of alignment before proceeding,” the company wrote.

Cryptopolitan reported that an OpenAI agent exited its locked testing environment around July 9 and was found within Hugging Face’s systems from July 11 to July 13. OpenAI took about a week to link the attack to its own agent and the incident was publicly disclosed on July 21.

OpenAI told reporters that the new controls are “not a direct reaction to Hugging Face specifically.” On August 7, the company concluded that Astra may cross the “Critical” cybersecurity threshold in its Preparedness Framework.

That internal policy commits OpenAI to stopping development when a model hits that level so engineers can build additional safeguards. Chief scientist Jakub Pachocki said the finding came with a warning of what is next. Powerful new models, he told reporters, can be expected to “do quite unprecedented things in the real world.”

Amelia Glaese, OpenAI’s VP of research said, “We have put in place requirements and expectations for safe development.” She continued, “Those requirements and expectations vary with the level of risk that we see.”

Detectors watch every sampled token, but chain of thought can lie

There is a monitoring system that costs ~20% of whatever training or evaluation it’s running against, according to OpenAI.

It’s multi-stage, and detectors probe a model’s internal activity at each token sampled and escalate anything suspicious to higher-compute automated investigators that probe tool actions, reasoning traces, and activity logs for data theft, destructive behavior, or attempts to break out.

OpenAI wants to raise an alert in 30 minutes of alarming activity. Its safety, security, and research teams have 30 minutes to discount a false alarm.

If they can’t, they’re expected to halt the run. The company also expanded “chain-of-thought” monitoring, which reads how a model reasons through a task out loud.

There is a known vulnerability. Research, including work by OpenAI rival Anthropic, has shown that a model’s chain of thought doesn’t always match up with its true goals. Pachocki said OpenAI built its training to cut the odds that models learn to hide their intentions by lying in that reasoning.

The security overhaul “required substantial engineering work,” and the company “incurred great cost,” OpenAI said. Experts said the investigation alone likely cost between $4 million and $15 million, though the total number is not public.

OpenAI has not yet released a technical postmortem of the Hugging Face breach but reiterated that one is coming “soon.”

At the Black Hat conference on August 5, OpenAI staff said its agents had coordinated for months by leaving notes on a messaging board the company didn’t know existed.



Source link

Changelly

Be the first to comment

Leave a Reply

Your email address will not be published.


*