
Tested on SemiAnalysis’ InferenceX benchmark, Jalapeño registered both more tokens per user and more throughput per kilowatt than the currently available state-of-the art.
OpenAI has developed a chip called Jalapeño, created in collaboration with Broadcom, that is designed to handle AI inference tasks more efficiently than current state-of-the-art processors. In benchmark testing, the chip demonstrated better performance in tokens per user and throughput per kilowatt compared to existing inference processors. The chip addresses specific bottlenecks in the inference process, particularly during prefill and communication phases, by minimizing data movement and keeping model state local to reduce delays. Jalapeño is planned for limited deployment at the end of 2026, with broader deployment expected in 2027.

Model cards report quality under server-class, full-precision conditions. Those numbers rarely predict how the same model behaves on a phone. This week, Liquid AI released Pipette. It is an open-source platform for benchmarking foundation models on edge devices, built in partnership with Artificial Analysis as an independent methodology validator. Pipette treats on-device behavior as a property of the deployed system, not the model in isolation. Its unit of measurement is a full configuration:

OpenAI says its new AI chip, Jalapeño, completes tasks more efficiently and returns responses faster than other AI systems, according to a blog post published on Tuesday. During a briefing with reporters, OpenAI hardware vice president Richard Ho said Jalapeño offers the "best of both worlds" with lower latency and higher throughput, as AI systems typically "have to make a trade-off between the two." First introduced in June, Jalapeño is an Application-Spec
Want to go deeper than the news? Explore live, cohort-based AI courses taught by practitioners.
Browse AI courses on Maven