
Jalapeño is a custom inference chip from OpenAI that delivers faster, more power-efficient AI inference, with higher throughput and lower latency for modern models.
OpenAI has announced results from Jalapeño, its first custom inference chip designed specifically for serving AI models. The chip achieves industry-leading performance by delivering more AI work per unit of power while also returning responses more quickly, addressing a traditional tradeoff where existing hardware systems usually have to choose between speed or efficiency. This matters because faster responses and greater efficiency can make AI more affordable and broadly available, particularly for interactive agents that require multiple sequential steps where delays compound. The chip was designed by considering hardware, memory, software, and systems together around real language-model workloads, with AI playing a direct role in the chip's development and programming.

IBM has rolled out the newest models in its family of open-weight large language models designed to be downloaded and self-hosted. The newly launched Granite 4.2 comes in 3B, 8B, and 30B parameter variants. Like previous versions, IBM is taking a decoder-only approach here. These new releases offer a 128,000-token context window natively. The 8B and 30B variants (not the 3B one) also go through an agentic reinforcement learning block; they were trained for expanded capabilities like using the t

IBM has released Granite 4.2, a family of open reasoning language models in 3B, 8B, and 30B parameter sizes. Unlike earlier Granite releases, which were instruction-following assistants, Granite 4.2 is built around explicit reasoning. Every model can emit a chain of thought before answering, and every model exposes a thinking / non-thinking switch plus a low-effort mode that spends a short reasoning budget on easy questions. The models are decoder-only dense transformers, pre-trained from scrat

Understanding the architecture and training methods behind enterprise-focused models helps explain their design tradeoffs compared to consumer alternatives.
Want to go deeper than the news? Explore live, cohort-based AI courses taught by practitioners.
Browse AI courses on Maven