
Vals AI is hoping to make AI benchmarking a more neutral and trustworthy resource in a world increasingly inundated by AI models.
Vals is a startup that tests artificial intelligence models to measure their real-world capabilities across industries like law, finance, and coding. The company matters because AI benchmarking has become an industry standard for validating model performance, but existing benchmarks are outdated and companies have learned to game them by training their models against publicly available tests. Vals differentiates itself by keeping its test materials private and evaluating whether models can complete complex tasks at human-level quality rather than just assessing general knowledge. The startup has grown rapidly after raising $40 million in Series A funding and now works with companies that pay to have their models evaluated, a revenue model the founder compares to students paying to take standardized tests.

MentalHealthBench is an expert-informed benchmark for evaluating helpful and safe AI responses across realistic mental health conversations.
Inconsistent benchmark testing has made it hard to compare AI systems fairly, creating a need for standardized evaluation methods.
Want to go deeper than the news? Explore live, cohort-based AI courses taught by practitioners.
Browse AI courses on Maven