
Model cards report quality under server-class, full-precision conditions. Those numbers rarely predict how the same model behaves on a phone. This week, Liquid AI released Pipette. It is an open-source platform for benchmarking foundation models on edge devices, built in partnership with Artificial Analysis as an independent methodology validator. Pipette treats on-device behavior as a property of the deployed system, not the model in isolation. Its unit of measurement is a full configuration:
Will Liquid AI's Pipette benchmarking suite appear on Hugging Face by September 5, 2026?
Resolves by Sep 5, 2026
Liquid AI released Pipette, an open-source benchmarking platform that measures how AI models perform on phones and other edge devices by testing complete deployment configurations rather than models in isolation. The tool matters because model performance reports from standard server conditions often do not predict how those same models actually behave on phones, and Pipette provides measurements across model, quantization method, runtime, device, and context length together. The platform launched with results from over 1,000 configurations spanning 30 or more models tested on MacBooks, iPhones, and Android devices, showing that seemingly similar models can retain vastly different performance as context lengths increase. The benchmarking suite is available as open-source infrastructure, a public results dashboard, and native mobile apps for developers deciding which models and settings to deploy on hardware they do not control.

OpenAI says its new AI chip, Jalapeño, completes tasks more efficiently and returns responses faster than other AI systems, according to a blog post published on Tuesday. During a briefing with reporters, OpenAI hardware vice president Richard Ho said Jalapeño offers the "best of both worlds" with lower latency and higher throughput, as AI systems typically "have to make a trade-off between the two." First introduced in June, Jalapeño is an Application-Spec

Tested on SemiAnalysis’ InferenceX benchmark, Jalapeño registered both more tokens per user and more throughput per kilowatt than the currently available state-of-the art.

Shrinking neural networks to use less memory and compute while maintaining or improving accuracy could make AI models cheaper to deploy widely.
Want to go deeper than the news? Explore live, cohort-based AI courses taught by practitioners.
Browse AI courses on Maven