
IBM has rolled out the newest models in its family of open-weight large language models designed to be downloaded and self-hosted. The newly launched Granite 4.2 comes in 3B, 8B, and 30B parameter variants. Like previous versions, IBM is taking a decoder-only approach here. These new releases offer a 128,000-token context window natively. The 8B and 30B variants (not the 3B one) also go through an agentic reinforcement learning block; they were trained for expanded capabilities like using the t
IBM has released new open-weight large language models called Granite 4.2 in three sizes that can be downloaded and run locally on personal computers rather than accessed through cloud services. The models are designed with features for enterprise use, including support for using external tools and a focus on reasoning capabilities that provide more rigorous responses, though often at the cost of slower speeds and higher computational demands. There has been growing interest in local models as cheaper alternatives to expensive cloud-based services, and organizations are increasingly using model routers to direct different tasks to appropriately sized models to balance performance, cost, and speed. IBM's approach emphasizes predictable deployments for enterprise customers rather than competing on speed or innovation, positioning these models as a practical solution within the expanding local AI deployment space.

IBM has released Granite 4.2, a family of open reasoning language models in 3B, 8B, and 30B parameter sizes. Unlike earlier Granite releases, which were instruction-following assistants, Granite 4.2 is built around explicit reasoning. Every model can emit a chain of thought before answering, and every model exposes a thinking / non-thinking switch plus a low-effort mode that spends a short reasoning budget on easy questions. The models are decoder-only dense transformers, pre-trained from scrat

Understanding the architecture and training methods behind enterprise-focused models helps explain their design tradeoffs compared to consumer alternatives.

Jalapeño is a custom inference chip from OpenAI that delivers faster, more power-efficient AI inference, with higher throughput and lower latency for modern models.
Want to go deeper than the news? Explore live, cohort-based AI courses taught by practitioners.
Browse AI courses on Maven