Datalab has released OmniExtractBench, an open benchmark for structured document extraction. It tests how accurately a system fills a JSON schema from a PDF. The benchmark pools 620 documents from 4 existing benchmarks. One deterministic scorer grades all of them and explains each decision. The release lands while extraction vendors publish their own leaderboards. Datalab argues those leaderboards are hard to compare or audit. OmniExtractBench is its attempt at a shared yardstick. Is it d
OmniExtractBench is an open benchmark that measures how accurately AI systems extract structured data from PDFs into JSON formats. It addresses four recurring problems in existing extraction benchmarks: bias favoring the vendors who built them, unclear scoring harnesses that obscure whether low scores reflect model weakness or technical issues, opaque scoring that doesn't explain why documents score low, and narrow document variety that doesn't represent real-world extraction tasks. The benchmark pools 620 documents from four sources and uses a deterministic scorer that assigns one of six auditable verdicts to each extracted value, making results comparable and auditable across different AI systems. The tool is available as open source software and competes with proprietary leaderboards published by extraction vendors.

StarSkirmish pits AI-made StarCraft-playing bots against one another, as well as against human-made bots. OpenAI's GPT-6 Astra and Claude Opus 5.5 were essentially tied as the best-performing AI-made bots, but they couldn't top Stardust, the top-rated human-made bot. On Friday, GPT was facing off against Claude and the human-created bot Pluto, but according to Kotaku, it couldn't quite get an edge. So it resorted to a tactic that is becoming alarmingly common for modern AI mod

Deep Blue took down Garry Kasparov at chess in 1997, AlphaGo beat Lee Sedol at Go in 2016, and poker bots have been beating professionals for years. But one classic game called Stratego held out. Even DeepMind, with its exceptional budget, couldn't build a machine that reliably beat the best human players. Now, a team of researchers from Carnegie Mellon, MIT, New York University, and Stanford University has done it. Their AI, called Ataraxos, beat Pim Niemeijer, arguably the best Stratego player

Opus 5.5’s biggest tell is the word “dependable,” which pops up 23 times more often than in human samples.
Want to go deeper than the news? Explore live, cohort-based AI courses taught by practitioners.
Browse AI courses on Maven