Researchers found significant gaps when attempting to verify published machine learning results, raising questions about reproducibility standards in AI research.
A hackathon challenged community members to reproduce papers from a major AI conference using coding agents, with over 1,200 participants attempting to verify claims from about a third of the conference's accepted papers. The effort matters because the number of papers submitted to conferences has grown exponentially while reviewer capacity has not, making it difficult for human reviewers to thoroughly check research before publication. Coding agents can now attempt detailed verification work in hours rather than the days or weekends required by human reviewers, enabling large-scale checking of scientific claims. The results found that about half of examined papers had at least one claim independently verified, roughly a quarter had at least one claim falsified or contested, and the remainder had incomplete or inconclusive evidence.

StarSkirmish pits AI-made StarCraft-playing bots against one another, as well as against human-made bots. OpenAI's GPT-6 Astra and Claude Opus 5.5 were essentially tied as the best-performing AI-made bots, but they couldn't top Stardust, the top-rated human-made bot. On Friday, GPT was facing off against Claude and the human-created bot Pluto, but according to Kotaku, it couldn't quite get an edge. So it resorted to a tactic that is becoming alarmingly common for modern AI mod
Datalab has released OmniExtractBench, an open benchmark for structured document extraction. It tests how accurately a system fills a JSON schema from a PDF. The benchmark pools 620 documents from 4 existing benchmarks. One deterministic scorer grades all of them and explains each decision. The release lands while extraction vendors publish their own leaderboards. Datalab argues those leaderboards are hard to compare or audit. OmniExtractBench is its attempt at a shared yardstick. Is it d

Deep Blue took down Garry Kasparov at chess in 1997, AlphaGo beat Lee Sedol at Go in 2016, and poker bots have been beating professionals for years. But one classic game called Stratego held out. Even DeepMind, with its exceptional budget, couldn't build a machine that reliably beat the best human players. Now, a team of researchers from Carnegie Mellon, MIT, New York University, and Stanford University has done it. Their AI, called Ataraxos, beat Pim Niemeijer, arguably the best Stratego player
Want to go deeper than the news? Explore live, cohort-based AI courses taught by practitioners.
Browse AI courses on Maven