Piloting the world's first double-blind AI evaluations
Double-blind AI evaluations are a new testing method where external evaluators assess AI models without the model provider seeing the test questions and without evaluators seeing the model's internal weights, using cryptographic safeguards to keep both sides private. This approach matters because AI models can artificially inflate their performance scores if they have already encountered test questions during training, a problem called benchmark contamination that undermines the trustworthiness of evaluation results. Historically, organizations faced a difficult choice: share test prompts and risk the model provider seeing them in advance, or share the model itself and expose proprietary information. By using cryptographic verification technology, double-blind evaluations eliminate this compromise and allow independent organizations to rigorously test advanced models while protecting both intellectual property and data security.

Google is rolling out an updated AI weather model that's supposed to be more accurate, especially when it comes to predicting rain and snowfall. In the announcement today, the company says it's now able to make forecasts with "unprecedented resolution" using its new WeatherNext 3 AI model. It can produce a global picture that's five times sharper than Google's previous model by learning from real-time weather observations, according to the company. "One of the main developments

Researchers question whether popular AI tests truly capture real-world language abilities or merely reflect memorization and training data overlap.

How do you benchmark a web search API when the thing being tested can read the answer key? A search agent has a fetch tool. If the gold labels sit in a public dataset, the agent can download them mid-evaluation and skip retrieval entirely. A similar problem arises when the answers are already encoded in the model’s parametric memory: a correct response no longer demonstrates that web search worked. Keenable’s answer is NEEDLE, a live open-source benchmark that rebuilds its query set from
Want to go deeper than the news? Explore live, cohort-based AI courses taught by practitioners.
Browse AI courses on Maven