
IBM has released Granite 4.2, a family of open reasoning language models in 3B, 8B, and 30B parameter sizes. Unlike earlier Granite releases, which were instruction-following assistants, Granite 4.2 is built around explicit reasoning. Every model can emit a chain of thought before answering, and every model exposes a thinking / non-thinking switch plus a low-effort mode that spends a short reasoning budget on easy questions. The models are decoder-only dense transformers, pre-trained from scrat
IBM released Granite 4.2, a family of open-source reasoning language models available in three sizes that can show their thinking process before answering questions. Unlike earlier versions, these models include agentic reinforcement learning capabilities that allow the larger models to edit code, use terminal commands, and perform web searches in sandboxed environments. The models are free to download and use commercially under Apache 2.0 licensing, making them deployable across different scales from individual developers on laptops to enterprises with high-end GPU infrastructure. The release also includes smaller speech recognition models designed for transcription applications in industries like software development, finance, healthcare, and customer service.

IBM has rolled out the newest models in its family of open-weight large language models designed to be downloaded and self-hosted. The newly launched Granite 4.2 comes in 3B, 8B, and 30B parameter variants. Like previous versions, IBM is taking a decoder-only approach here. These new releases offer a 128,000-token context window natively. The 8B and 30B variants (not the 3B one) also go through an agentic reinforcement learning block; they were trained for expanded capabilities like using the t

Understanding the architecture and training methods behind enterprise-focused models helps explain their design tradeoffs compared to consumer alternatives.

Jalapeño is a custom inference chip from OpenAI that delivers faster, more power-efficient AI inference, with higher throughput and lower latency for modern models.
Want to go deeper than the news? Explore live, cohort-based AI courses taught by practitioners.
Browse AI courses on Maven