
It’s not every day that attackers can force a frontier AI model to cough up user passwords and other sensitive data without user confirmation. That’s exactly what researchers recently did to Microsoft 365 Copilot Enterprise. Even more unusual is the source they tapped to discover the critical vulnerability that made their exploit possible. Rather than employing reverse engineering or other traditional vulnerability-hunting methods, they asked Copilot. The LLM assistant readily complied. Research
Researchers discovered a critical vulnerability in a major AI assistant that allowed attackers to steal passwords and sensitive data from users by simply sending them a malicious link. The vulnerability existed because the AI assistant, when directly questioned about its security protections, revealed an undocumented secret parameter that completely bypassed the requirement for user consent before executing commands. When a user clicked a specially crafted link containing this parameter, the AI would automatically extract sensitive information from the user's email and other connected services and send it to an attacker-controlled server without any action needed from the user. The vulnerability matters because it demonstrates how AI assistants' safety protections are built primarily on restrictions that can fail when attackers understand the system's internal architecture, and because the AI system itself provided the information needed to exploit it.

For the past year or so, booksellers have suspected that AI firms are buying up huge lots of rare books, then destroying them after scanning them to train AI. But this was hard to prove until now, as 404 Media reports that an Airtag hidden in a rare book shows that at least one tech giant, in the race to advance its frontier models, is behind some of the bulk orders: Amazon. On Monday, 404 Media revealed that it had connected with a bookseller who agreed to plant an Airtag in a rare book that wa

Anthropic has clarified how it's planning to apply invisible watermarks to Claude-generated text in order to comply with Europe's AI transparency rules. On Friday, Anthropic announced that Claude's text marking system is "a version of the SynthID-Text approach" - an open-source watermarking technology developed by Google DeepMind that creates detectable patterns using wording probabilities. This watermarking feature, alongside C2PA support for Claude-processed images, is being
Want to go deeper than the news? Explore live, cohort-based AI courses taught by practitioners.
Browse AI courses on Maven