Inconsistent benchmark testing has made it hard to compare AI systems fairly, creating a need for standardized evaluation methods.
The UK AI Security Institute is using EvalEval's shared infrastructure to publicly release evaluation results for AI models in a standardized format, making those results more reproducible and verifiable. Evaluation results are currently reported across many different formats and platforms without sufficient information to reproduce them, making it difficult for researchers to verify findings or compare results across studies. The two organizations are collaborating to adopt a shared schema called Every Eval Ever and an open platform called Evaluation Cards that combine benchmark data, evaluation-run information, and model metadata into a common structure. This effort aims to improve the quality and reliability of AI evaluation science by creating transparent reference points that help researchers understand how different evaluation setups influence reported model performance.

MentalHealthBench is an expert-informed benchmark for evaluating helpful and safe AI responses across realistic mental health conversations.

Vals AI is hoping to make AI benchmarking a more neutral and trustworthy resource in a world increasingly inundated by AI models.
Want to go deeper than the news? Explore live, cohort-based AI courses taught by practitioners.
Browse AI courses on Maven