
In this tutorial, we use Anthropic’s claude-protein-binder-design dataset, which contains 1,440 AI-designed miniprotein binders tested against 16 targets. Because the release includes both computational predictions and real wet-lab results from two independent labs, we can go beyond simply studying the designs. We evaluate how well structure predictors identify successful binders, whether combining predictions improves performance, how rankings translate into practical testing budgets, and how
This tutorial evaluates AI-designed miniprotein binders by comparing computational predictions with real experimental results from two independent labs. The dataset contains over 1,400 AI-designed proteins tested against multiple targets, allowing researchers to assess how well structure predictors identify successful binders and whether computational signals can reliably predict whether designs will work in practice. The analysis examines hit rates across different design models, campaigns, and targets, revealing that target choice has a larger effect on success rates than the choice of design method. Understanding this gap between computational design and wet-lab validation matters because it shows which aspects of AI protein design translate reliably into functional molecules.

Google is rolling out an updated AI weather model that's supposed to be more accurate, especially when it comes to predicting rain and snowfall. In the announcement today, the company says it's now able to make forecasts with "unprecedented resolution" using its new WeatherNext 3 AI model. It can produce a global picture that's five times sharper than Google's previous model by learning from real-time weather observations, according to the company. "One of the main developments

Researchers question whether popular AI tests truly capture real-world language abilities or merely reflect memorization and training data overlap.

How do you benchmark a web search API when the thing being tested can read the answer key? A search agent has a fetch tool. If the gold labels sit in a public dataset, the agent can download them mid-evaluation and skip retrieval entirely. A similar problem arises when the answers are already encoded in the model’s parametric memory: a correct response no longer demonstrates that web search worked. Keenable’s answer is NEEDLE, a live open-source benchmark that rebuilds its query set from
Want to go deeper than the news? Explore live, cohort-based AI courses taught by practitioners.
Browse AI courses on Maven