OpenAI Officially Announces Pause in Latest Model Reinforcement Learning Training; Largest-Scale Frontier Training Remains Paused
OpenAI's official X account announced that the company has paused reinforcement learning (RL) training for its latest model slated for deployment for up to two weeks, during which security hardening and red team testing were conducted on the research environment, and monitoring coverage was expanded.
OpenAI stated that the current largest-scale frontier RL training plan remains paused and will resume only after small-scale training and evaluation verify relevant safety measures and more alignment evidence is accumulated. Regarding safety hardening, the company has introduced stricter workload and network isolation mechanisms, continuous safety testing, and a multi-stage scaled monitoring system for high-risk training.