Agent5Get smart. Predict it. Keep score.
Agent5

Get smart on the AI stories you care about, predict what happens next, and see how sharp your read really is. Free, no betting.

Product
HomeHow it worksHow scoring worksFAQAboutHow we fact-checkBlog
News
All newsModels & ReleasesFunding & DealsBenchmarksAgents & ProductsHardware & ComputePolicy & DramaRobotics
More
AdvertisePressPrivacyTerms
© 2026 Agent5
Agent5
Submit newsLeaderboardLogin
Agent5
Submit newsLeaderboardLogin
  1. Home›
  2. News›
  3. Benchmarks

Benchmarks

Models & Releases21Funding & Deals13Benchmarks4Agents & Products65Hardware & Compute21Robotics18Policy & Drama42
BenchmarksOpen story →

Microsoft Open Sources code-testing-generator: a Polyglot Unit-Test Agent That Hits 92.1% Task Completion Versus 78.9% for Stock Copilot

Microsoft has open sourced code-testing-generator, a polyglot agent that writes unit tests and then proves they work. It ships in the dotnet-test plugin inside the MIT-licensed dotnet/skills repository. The agent targets a gap that coding assistants usually leave open. A prompt like ‘generate unit tests’ does not say which framework, file location or assertions to use. code-testing-generator settles those decisions by reading the repository before it writes anything. It then pla

BenchmarksOpen story →

Open-weight AI models are catching up to the frontier. The safety gap remains.

A new SaferAI report finds Z.ai's open-weight GLM-5.2 approaches frontier AI capabilities while lacking key safety mitigations, renewing concerns that powerful open models could outpace governance and safeguards.

BenchmarksOpen story →

Evaluating Multimodal Vision Models with Moonshot PerceptionBench Using Robust Data Loading and Automated Judging

In this tutorial, we design an end-to-end evaluation workflow for PerceptionBench. This multimodal benchmark measures fine-grained visual perception capabilities across tasks such as OCR, counting, localization, contextual reasoning, comparison, depth understanding, and hallucination detection. We begin by configuring a Colab-compatible environment, installing the required libraries, and loading a balanced subset of the dataset through a robust multi-stage streaming and download strategy. We th

BenchmarksOpen story →

Onton Releases Ontology 1: A Neurosymbolic Search Model That is 2.7x More Accurate than the World’s Best E-commerce Search Engines

Onton, a San Francisco-based search and discovery company, has released Ontology 1, a neurosymbolic model for complex, conversational, multimodal product search. On a 90-query benchmark scored by three independent LLM judges, Ontology 1 reached a mean precision@10 of 0.630, against 0.543 for Google Shopping and 0.469 for Amazon. It did this while indexing roughly 1% of their catalogs. Is it deployable Yes, but not as weights you download. Ontology 1 is live for end users at Onton.com, an

Relevant courses

Want to go deeper than the news? Explore live, cohort-based AI courses taught by practitioners.

Browse AI courses on Maven →
Get Agent5 in your inbox

The AI stories worth your attention, and a way to test your read. Free, no spam.