
MentalHealthBench is an expert-informed benchmark for evaluating helpful and safe AI responses across realistic mental health conversations.
MentalHealthBench is an open benchmark created with more than 80 licensed mental health experts from 22 countries to measure how AI systems respond in realistic mental health conversations. The benchmark assesses model capabilities across key behaviors like safety, seeking context, preserving user agency, and providing actionable guidance, covering a full spectrum of situations from everyday stress to mental health emergencies. It matters because most previous AI evaluations in mental health have focused primarily on emergency scenarios, leaving a gap in understanding how models perform across the full range of conversations people have about well-being and life advice. The benchmark uses synthetic conversations that reflect real-world usage patterns and expert-written criteria with weighted scores to evaluate whether AI responses align with clinical guidance.
Inconsistent benchmark testing has made it hard to compare AI systems fairly, creating a need for standardized evaluation methods.

Vals AI is hoping to make AI benchmarking a more neutral and trustworthy resource in a world increasingly inundated by AI models.
Want to go deeper than the news? Explore live, cohort-based AI courses taught by practitioners.
Browse AI courses on Maven