
Generative AI
Researchers introduced a framework called knowledge profiling that distinguishes between two different reasons why language models get facts wrong: either they never learned the fact in the first place, or they learned it but cannot retrieve it when needed. The distinction matters because encoding failures and recall failures require different solutions, with encoding problems addressed through scaling and data expansion, while recall problems might benefit from post-training and inference-time methods. Testing frontier language models revealed that nearly all facts are encoded, yet these models frequently fail to recall encoded facts, suggesting that the main bottleneck for factual accuracy is shifting from learning facts to accessing them. The researchers created WikiProfile, a benchmark of 2,150 Wikipedia-derived facts with multiple question variations, to measure both encoding and recall across different language models.
As open-source AI models proliferate, tracking their capabilities against closed alternatives helps developers choose which tools suit their needs and budgets.
Researchers found significant gaps when attempting to verify published machine learning results, raising questions about reproducibility standards in AI research.

Object removal models have improved faster than the metrics used to judge them. Diffusion erasers now reconstruct shadows, reflections and occluded structure convincingly, yet PSNR, SSIM, LPIPS, ReMOVE and CFD frequently rank their outputs the wrong way. The root cause is structural: erasure is an ill-posed, one-to-many task, so no single ground truth exists to compare against. A team from MiLM Plus, Xiaomi Inc. has released PROVE (Perceptual RemOVal cohErence), accepted at ACM MM 2026, to clos
Want to go deeper than the news? Explore live, cohort-based AI courses taught by practitioners.
Browse AI courses on Maven