Expanding speech recognition beyond wealthy English-speaking markets could unlock AI applications for billions of underserved users worldwide.
The Open ASR Leaderboard is adding evaluation datasets for Hindi and Indian English, marking the first languages from the Global South on a multilingual benchmark that previously contained only European languages. ASR (automatic speech recognition) leaderboards measure how well speech-to-text models perform, but existing benchmarks record only what was said, not details about who said it, which means performance disparities across different populations remain invisible. The new datasets were deliberately built to vary across nine dimensions including geography, age, gender, vocabulary, devices, and acoustic environments, so that models cannot score well on average while performing poorly for specific groups. This approach addresses research showing that ASR error rates are not evenly distributed across people, with documented racial, gender, age, and accent-based disparities in commercial systems.

Google is rolling out an updated AI weather model that's supposed to be more accurate, especially when it comes to predicting rain and snowfall. In the announcement today, the company says it's now able to make forecasts with "unprecedented resolution" using its new WeatherNext 3 AI model. It can produce a global picture that's five times sharper than Google's previous model by learning from real-time weather observations, according to the company. "One of the main developments

Researchers question whether popular AI tests truly capture real-world language abilities or merely reflect memorization and training data overlap.

How do you benchmark a web search API when the thing being tested can read the answer key? A search agent has a fetch tool. If the gold labels sit in a public dataset, the agent can download them mid-evaluation and skip retrieval entirely. A similar problem arises when the answers are already encoded in the model’s parametric memory: a correct response no longer demonstrates that web search worked. Keenable’s answer is NEEDLE, a live open-source benchmark that rebuilds its query set from
Want to go deeper than the news? Explore live, cohort-based AI courses taught by practitioners.
Browse AI courses on Maven