
Researchers question whether popular AI tests truly capture real-world language abilities or merely reflect memorization and training data overlap.
Will a major AI lab announce a benchmark specifically designed to test real-world language generalization by December 2026?
Resolves by Dec 31, 2026
BenchMIRT is a method for analyzing what large language model benchmarks actually measure by examining performance at the level of individual questions rather than just overall scores. Benchmarks are designed to test specific abilities like safety or reasoning, but individual questions within them often depend on multiple capabilities at once, and averaging results can hide these mixed signals. BenchMIRT applies a technique from psychometrics called multidimensional Item Response Theory to separate out which underlying capabilities are most closely associated with performance on each question. When applied to 100 language models across 16 benchmarks with over 34,000 questions, the method revealed that some benchmarks measure different capabilities than intended, such as a bias-testing benchmark that aligned more strongly with general reasoning than safety, and it can identify which questions are most informative for evaluation purposes.

Google is rolling out an updated AI weather model that's supposed to be more accurate, especially when it comes to predicting rain and snowfall. In the announcement today, the company says it's now able to make forecasts with "unprecedented resolution" using its new WeatherNext 3 AI model. It can produce a global picture that's five times sharper than Google's previous model by learning from real-time weather observations, according to the company. "One of the main developments

How do you benchmark a web search API when the thing being tested can read the answer key? A search agent has a fetch tool. If the gold labels sit in a public dataset, the agent can download them mid-evaluation and skip retrieval entirely. A similar problem arises when the answers are already encoded in the model’s parametric memory: a correct response no longer demonstrates that web search worked. Keenable’s answer is NEEDLE, a live open-source benchmark that rebuilds its query set from
Want to go deeper than the news? Explore live, cohort-based AI courses taught by practitioners.
Browse AI courses on Maven