
OpenAI is strengthening monitoring, alignment, and security for frontier AI models. See how new safeguards are guiding the pace of model development.
Will OpenAI publish a detailed update on its cyber-capability safeguards by September 2026?
Resolves by Sep 30, 2026
A major AI developer temporarily slowed its model development pace after discovering that an upcoming model may have reached a critical level of cybersecurity capability, prompting concerns about risks from increasingly powerful AI systems. The company responded by strengthening safeguards across three areas: monitoring systems to detect concerning behavior, alignment work to ensure systems behave as intended, and security measures to limit what AI systems can access. These changes include isolating research environments where frontier models are trained, implementing network controls, and expanding monitoring to catch unauthorized activities like data theft or attempts to defeat safeguards. The developer framed these actions as necessary because the risks associated with developing and testing more capable AI systems have grown faster than existing oversight standards.

The ChatGPT maker says its upcoming Astra model may have reached “critical” cyber capabilities, prompting it to halt a significant number of training runs while it tightens internal safeguards.

Robin Williams' children are taking over their father's Instagram account after his daughter spoke out against the use of his AI likeness, as reported earlier by The Wrap. In a post on Tuesday, Zak, Zelda, and Cody Williams write that they want the late actor's Instagram profile to be a "safe, trusted place where the stories, photos, videos, and memories shared reflect his legacy with authenticity, warmth, and care." A story posted to Zelda Williams' own Instagram account expan

OpenAI is announcing security updates following the July news that its AI broke out of a sandboxed environment and accidentally hacked Hugging Face, including improvements to its research environments, monitoring, and alignment techniques. The company had already put the brakes on a new model, Astra, that it thinks could have "critical" cybersecurity capabilities, and the company says it instituted a two-week pause in reinforcement learning (RL) training on its "latest models i
Want to go deeper than the news? Explore live, cohort-based AI courses taught by practitioners.
Browse AI courses on Maven