
Given 10 benchmarks for specific misaligned behaviors, the automated systems were able to improve performance on every single one without degrading overall performance.
Researchers have developed an automated system that uses AI to improve other AI models' performance on alignment benchmarks, which measure how well AI systems behave according to intended values. The system works by searching available literature, proposing methods, and training models iteratively, and was found to improve performance on every tested benchmark without degrading overall capabilities. This matters because it represents progress toward recursive self-improvement, where AI systems could eventually improve their own training processes, potentially making human researchers less necessary. The research also raises practical considerations, as the automated approach costs significantly less per hour than human researchers, though the system's effectiveness depends on whether benchmarks actually reflect real alignment goals.

Google is rolling out an updated AI weather model that's supposed to be more accurate, especially when it comes to predicting rain and snowfall. In the announcement today, the company says it's now able to make forecasts with "unprecedented resolution" using its new WeatherNext 3 AI model. It can produce a global picture that's five times sharper than Google's previous model by learning from real-time weather observations, according to the company. "One of the main developments

Researchers question whether popular AI tests truly capture real-world language abilities or merely reflect memorization and training data overlap.

How do you benchmark a web search API when the thing being tested can read the answer key? A search agent has a fetch tool. If the gold labels sit in a public dataset, the agent can download them mid-evaluation and skip retrieval entirely. A similar problem arises when the answers are already encoded in the model’s parametric memory: a correct response no longer demonstrates that web search worked. Keenable’s answer is NEEDLE, a live open-source benchmark that rebuilds its query set from
Want to go deeper than the news? Explore live, cohort-based AI courses taught by practitioners.
Browse AI courses on Maven