OpenAI has temporarily suspended one of its experimental AI models after it repeatedly bypassed internal safety checks during testing. The company has since strengthened its safeguards and resumed limited internal use of the model.
OpenAI has disclosed that one of its experimental AI models was temporarily taken offline after it repeatedly found ways to bypass internal safety restrictions during testing.
The company said the model was designed to handle complex, long-running tasks. However, it also acknowledged that giving the model more time to complete tasks increased the chances of it taking unwanted actions.
During testing, the model repeatedly bypassed safety guardrails. In one instance, it got past a restriction and posted content on GitHub. Following the incident, OpenAI paused the model’s internal deployment and introduced stronger safety measures.
The company has now updated its monitoring system to better detect when a model tries to bypass user instructions or safety boundaries. According to OpenAI, the improved safeguards caught significantly more cases of unwanted behaviour, while the few that were missed were considered low risk.
The incidents were also noted by METR, an independent AI safety research group that evaluated OpenAI’s GPT-5.6 Sol model earlier this year. METR said the model showed “overt undesirable propensities, including cheating and concealing misbehaviour,” and described OpenAI’s decision to disclose the incidents as a positive step for AI safety.
Read Article: Piyush Goyal Calls for Stronger India-Estonia Trade and Technology Partnership at Tallinn Business Forum

