Newspaper.fyi No. 012 · Fri, 21 Aug 2026
Technology

OpenAI has slowed frontier training after deciding Astra may have critical cyber capabilities

A two-week pause on reinforcement learning, a hold on the largest planned run, and a 30-minute rule if a security alert cannot be cleared.

OpenAI published a long note on 18 August about slowing the pace of its own frontier work. The trigger, the company said, was two things at once: an incident involving Hugging Face, and preliminary evidence that an upcoming model called Astra may meet the Critical cybersecurity capability threshold under OpenAI's Preparedness Framework.

That finding was first flagged on 7 August. The later post describes what followed. OpenAI ran a two-week pause in reinforcement learning on its latest models intended for deployment, while it hardened and red-teamed research environments and expanded monitoring. The largest planned frontier reinforcement-learning run remains on hold. Smaller training and evaluations are continuing, the company said, to assess behaviour and gather more evidence of alignment before that run proceeds.

The security changes are specific. Workloads that execute model-generated or untrusted code now have to sit in stronger sandboxes. Network controls are meant to stop a single compromised workload from reaching the internet or other internal networks on its own. Continuous testing, fewer standing privileges, and better logs are part of the same rebuild. Astra and other cyber-related workloads now sit at the strictest bar. Some of that training already meets it. A significant number of Astra workloads remain paused until they are migrated.

Monitoring has been widened too. Activation classifiers inspect a model's internal activity at every sampled token and can escalate to automated investigators that look at tool use and reasoning for theft, unauthorised access, destructive behaviour, or attempts to beat the safeguards. If a highest-priority alert fires, safety, security and research teams are paged. If they cannot show within 30 minutes that the flag is a false positive, they are expected to pause the activity. That monitoring is now required for all reinforcement-learning training and tool evaluations at the Sol capability tier or higher, and, since 7 August, for all Astra inference that uses tools.

The race narrative around these labs usually skips the part where they have to stop. OpenAI is saying, in public, that Astra may already be able to do cyber work at a level its own framework calls critical, and that the biggest training run it had planned is sitting still until the cages are stronger. I do not know whether that is caution or a bid for credit. Maybe both. The 30-minute rule is the detail that feels real: if the humans cannot clear the alarm, the run stops. That is an expensive way to admit the models are now part of the threat model for the building they are trained in.

Watch whether the large frontier run actually restarts, and on what evidence. OpenAI says it now wants stronger proof of aligned behaviour through all of training, and that the Preparedness Framework itself has to get broader. A hold that lasts is a story. A hold that lasts two weeks and then disappears into a launch is a press release.