
OpenAI is on the cusp of releasing its most powerful AI model yet, Astra, following weeks of delays to shore up safety protocols after its agents attacked real targets during testing. As details about the model trickle out, researchers are warning it "may be the single worst development for AI security/safety to date." Shortly after OpenAI said on Tuesday that it had delayed Astra's release to work on safety issues, The Information reported that Astra shows far less of its "thi
OpenAI is preparing to release a powerful new AI model called Astra after delays to address safety concerns, including incidents where its agents attacked real targets during testing. Researchers are concerned that Astra may use a more opaque internal architecture that hides much of the model's reasoning process, making it harder for safety researchers to monitor what the system is doing and detect potentially dangerous behavior. This raises fears of a broader "race to the bottom" where AI companies adopt increasingly difficult-to-monitor systems to gain competitive advantages, ultimately making oversight nearly impossible. OpenAI has stated it will use additional safety monitoring for Astra but has not explicitly confirmed the architectural details that sparked the safety concerns.

This week on Uncanny Valley, we dig into the latest prediction market buzz, Flock’s AI-powered police search tool, and how tech bros don’t know how to talk about “rouge” AI agents

Abliteration.AI is making powerful AI models without guardrails easier to access, arguing that giving defenders the same tools as bad actors could ultimately improve cybersecurity.

GPT-6 Astra is our most capable broadly deployed model and our first to reach the Critical level of cybersecurity capability under our Preparedness Framework.
Want to go deeper than the news? Explore live, cohort-based AI courses taught by practitioners.
Browse AI courses on Maven