IBM’s new Granite 4.2 models add reasoning and stay dense
Summary
On Tuesday, IBM launched the latest family of its open-weight Granite large language models (LLMs). Weighing in at 3 billion, The post IBM’s new Granite 4.2 models add reasoning and stay dense appeared first on The New Stack .
Original Text
On Tuesday, IBM launched the latest family of its open-weight Granite large language models (LLMs). Weighing in at 3 billion, 8 billion, and 30 billion parameters, IBM is taking a very different approach to model building from some of the competition here, with dense, decoder-only reasoning models that it pre-trained from scratch.
Many recent models have moved from all-attention Transformers toward hybrid Mamba/attention architectures, including Nvidia’s Nemotron 3 family. But IBM tried that with its Granite 4.0 models. That generation included conventional dense, dense-hybrid, and hybrid MoE models.
With Granite 4.1, IBM returned the main family to an all-attention, dense Transformer architecture. At the time, IBM said that these new models outperformed the older generation, “while using a simpler — and therefore more flexible — architecture for fine-tuning for downstream tasks.”
Reasoning with Granite
IBM describes the 4.2 family as a “reasoning-focused release.” Models can run in thinking and non-thinking modes, but there is also a low-effort mode that only spends a low number of reasoning tokens for answering easy questions.
When the company launched the Granite 4.1 models, IBM still argued that reasoning models weren’t efficient enough, so “turning to less expensive, non-reasoning models with similar benchmark performance for select tasks like instruction following and tool calling makes sense for enterprise users.” At this point, though, the team clearly believes that reasoning is a necessary feature — even as it remains optional for the 4.2 models.
Unlike some other models in its class, including Qwen-3.8 27B, Muse Glimmer 30B, and Google’s Gemma 4 31B, it’s a text-only model, though. The others are all multi-modal, though IBM does offer Granite Vision 4.1 4B, for example, and there’s always a chance we’ll get a 4.2 version of this model, too.
It’s worth noting that IBM also launched two new speech recognition models in the Granite Speech family on Tuesday as well.
Training Granite
The Apache 2.0-licensed models were pre-trained on 15 trillion tokens and in five phases, including a long-context training phase that now brings the family’s context window to 512,000 tokens (although the released configuration natively supports 128K).
IBM notes that the model’s training set also included 1 trillion tokens of synthetic code, generated by IBM’s CodeAlchemy pipeline (though the models’ overall coding performance remains average).
For the most part, all of the models share this same pipeline, but the 8B and 30B models also went through an additional agentic reinforcement learning step to allow them to call tools, edit and run code, work in the terminal, and search the web. “Combined with reinforcement learning from human feedback (RLHF) alignment, this approach produces models better equipped for complex, multi-step agentic work,” IBM writes.
The 3B model also supports tool calling. Just don’t expect too much from it.
Benchmarks? It’s fine.
In terms of benchmarks, the Granite 4.2 models aren’t breaking any new ground, but it’s worth noting that the small 8B model often gets very close to the results of the larger 30B model — and it’s easy to run on virtually any modern Mac and even some relatively lower-end Nvidia RTX GPUs.
Qwen 3.8 27B, especially, beats the Granite models across the board, especially when it comes to coding, where the IBM models deliver inconsistent results overall.
As always, though, benchmarks only tell part of the story. For the right kind of usecase, frontier models are often overkill and IBM argues that the main argument for Granite is its performance in high-throughput agentic tasks, for example.
“This release extends the Granite family with a clear goal: helping enterprises build agents that can reason, act, and adapt during real-life workflows,” IBM writes in its announcement. “Now that AI systems are being asked to carry out tasks in the real world, our expectations have risen. It’s no longer enough to answer clearly and concisely. An AI system must be able to plan, call applications, and execute complex tasks in a reliable and consistent way — while staying light enough to actually use without breaking the bank.”
The post IBM’s new Granite 4.2 models add reasoning and stay dense appeared first on The New Stack.
Lotu Radar provides attributed news summaries and links to the original publisher. Full reporting and copyright remain with the source.