
Learn how OpenAI disrupted a campaign to extract protected model reasoning and is strengthening defenses against adversarial distillation.
OpenAI identified and disrupted a coordinated campaign in which operators systematically extracted protected reasoning from OpenAI's models through manipulated interactions, rather than through direct hacking or database breaches. Protected reasoning is the model's internal working process, and extracting it can reveal information hidden from final answers and help others reproduce the model's capabilities. This type of attack, called adversarial distillation, poses safety and national security risks because extracted reasoning could be used to train other models without the original safeguards, and at scale it can accelerate the transfer of advanced capabilities without equal investment in safety measures. OpenAI responded by banning fraudulent accounts, strengthening technical controls, and sharing findings with industry partners and government channels so other developers can defend against similar attacks.

Well, if AI said it, it must be true. | Bloomberg via Getty Images New Jersey's lieutenant governor Dale Caldwell was forced to resign on September 25th after an investigation found he had sexually harassed a staffer and repeatedly violated ethics rules. The now-former Lt. governor has been making the media rounds trying to clear his name. But he took a particularly odd tactic during an interview on NJ PBS. Caldwell claims he's being unfairly targeted, and multiple AI agents ba

This new task force is Trump's latest response to the debate over AI safety.

By his own admission, David Robinson is “something of a cliché”: an employee at a leading AI company who issues a dire warning while resigning from their job.
Want to go deeper than the news? Explore live, cohort-based AI courses taught by practitioners.
Browse AI courses on Maven