Researchers found significant gaps when attempting to verify published machine learning results, raising questions about reproducibility standards in AI research.
A hackathon challenged community members to reproduce papers from a major AI conference using coding agents, with over 1,200 participants attempting to verify claims from about a third of the conference's accepted papers. The effort matters because the number of papers submitted to conferences has grown exponentially while reviewer capacity has not, making it difficult for human reviewers to thoroughly check research before publication. Coding agents can now attempt detailed verification work in hours rather than the days or weekends required by human reviewers, enabling large-scale checking of scientific claims. The results found that about half of examined papers had at least one claim independently verified, roughly a quarter had at least one claim falsified or contested, and the remainder had incomplete or inconclusive evidence.
As open-source AI models proliferate, tracking their capabilities against closed alternatives helps developers choose which tools suit their needs and budgets.

Generative AI

Object removal models have improved faster than the metrics used to judge them. Diffusion erasers now reconstruct shadows, reflections and occluded structure convincingly, yet PSNR, SSIM, LPIPS, ReMOVE and CFD frequently rank their outputs the wrong way. The root cause is structural: erasure is an ill-posed, one-to-many task, so no single ground truth exists to compare against. A team from MiLM Plus, Xiaomi Inc. has released PROVE (Perceptual RemOVal cohErence), accepted at ACM MM 2026, to clos
Want to go deeper than the news? Explore live, cohort-based AI courses taught by practitioners.
Browse AI courses on Maven