
In this tutorial, we design an end-to-end evaluation workflow for PerceptionBench. This multimodal benchmark measures fine-grained visual perception capabilities across tasks such as OCR, counting, localization, contextual reasoning, comparison, depth understanding, and hallucination detection. We begin by configuring a Colab-compatible environment, installing the required libraries, and loading a balanced subset of the dataset through a robust multi-stage streaming and download strategy. We th
PerceptionBench is a multimodal benchmark that evaluates vision models across tasks including optical character recognition, counting, localization, and reasoning about images. The tutorial describes building an end-to-end evaluation workflow that loads the benchmark dataset through a multi-stage streaming strategy, processes images, and runs models through different backends while measuring performance using both rule-based and language model-assisted judging. This matters because it provides researchers with a structured method to assess how well vision models perform on fine-grained visual perception capabilities and compare results against a leaderboard. The workflow supports multiple evaluation approaches, from simple baselines to cloud-based APIs and local open-source models, making it adaptable to different computational environments and needs.

Microsoft has open sourced code-testing-generator, a polyglot agent that writes unit tests and then proves they work. It ships in the dotnet-test plugin inside the MIT-licensed dotnet/skills repository. The agent targets a gap that coding assistants usually leave open. A prompt like ‘generate unit tests’ does not say which framework, file location or assertions to use. code-testing-generator settles those decisions by reading the repository before it writes anything. It then pla

A new SaferAI report finds Z.ai's open-weight GLM-5.2 approaches frontier AI capabilities while lacking key safety mitigations, renewing concerns that powerful open models could outpace governance and safeguards.

Onton, a San Francisco-based search and discovery company, has released Ontology 1, a neurosymbolic model for complex, conversational, multimodal product search. On a 90-query benchmark scored by three independent LLM judges, Ontology 1 reached a mean precision@10 of 0.630, against 0.543 for Google Shopping and 0.469 for Amazon. It did this while indexing roughly 1% of their catalogs. Is it deployable Yes, but not as weights you download. Ontology 1 is live for end users at Onton.com, an
Want to go deeper than the news? Explore live, cohort-based AI courses taught by practitioners.
Browse AI courses on Maven