source: Hugging Face Blog: Granite 4.2 LLMs: How They're Built

level: technical

ibm released granite 4.2, a family of dense decoder-only reasoning llms in three sizes: 3b, 8b, and 30b parameters. each model is pre-trained from scratch on roughly 15 trillion tokens using a five-phase strategy that extends the context window to 512k tokens. after supervised fine-tuning on chain-of-thought, reasoning, and agentic-trajectory data, the models go through a multi-stage reinforcement learning pipeline. the 8b and 30b models also receive agentic rl, learning to use tools inside real sandboxed environments.

the training pipeline is detailed and resource-intensive. supervised fine-tuning uses about 7.2 million samples, or 100b tokens, with 31.6% agentic data covering software engineering, tool calling, and terminal use. reinforcement learning runs as a chain of stages: math, code, science, instruction following, tool use, then software engineering, terminal, and web search. each stage uses asynchronous grpo with group-relative advantages. for the 30b model, the rlvr stage processes 256 prompts with 16 responses each per step, a batch of 4,096 examples.

granite 4.2 models include a thinking and non-thinking switch, a low-effort mode for easy questions, and native tool calling. they are released under the apache 2.0 license and served through openai-compatible endpoints. the 3b model skips the agentic rl block, while the 8b and 30b models complete the full ladder. this staged approach lets smaller models stay efficient while larger ones gain agentic abilities. the design shows a practical trade-off between capability and compute cost.

why it matters: granite 4.2 shows how to build reasoning llms with verifiable rewards and agentic training, giving ai teams a reproducible recipe for smaller, capable models.


source: Hugging Face Blog: Granite 4.2 LLMs: How They're Built