
Generative AI
Researchers introduced a framework called knowledge profiling that distinguishes between two different reasons why language models get facts wrong: either they never learned the fact in the first place, or they learned it but cannot retrieve it when needed. The distinction matters because encoding failures and recall failures require different solutions, with encoding problems addressed through scaling and data expansion, while recall problems might benefit from post-training and inference-time methods. Testing frontier language models revealed that nearly all facts are encoded, yet these models frequently fail to recall encoded facts, suggesting that the main bottleneck for factual accuracy is shifting from learning facts to accessing them. The researchers created WikiProfile, a benchmark of 2,150 Wikipedia-derived facts with multiple question variations, to measure both encoding and recall across different language models.

Deep Blue took down Garry Kasparov at chess in 1997, AlphaGo beat Lee Sedol at Go in 2016, and poker bots have been beating professionals for years. But one classic game called Stratego held out. Even DeepMind, with its exceptional budget, couldn't build a machine that reliably beat the best human players. Now, a team of researchers from Carnegie Mellon, MIT, New York University, and Stanford University has done it. Their AI, called Ataraxos, beat Pim Niemeijer, arguably the best Stratego player

Opus 5.5’s biggest tell is the word “dependable,” which pops up 23 times more often than in human samples.
Comparing speech synthesis systems across languages and voices has lacked standardized metrics, making it harder to track progress in this rapidly advancing field.
Want to go deeper than the news? Explore live, cohort-based AI courses taught by practitioners.
Browse AI courses on Maven