Skip to content
Main Site News Console

Granite 4.2 Large Language Models: How They’re Built

· Hugging Face Translated
教程模型卡

A Technical Deep Dive into the Granite 4.2 Reasoning Model Series

Authors: The Granite Team, IBM

Overview: Granite 4.2 is our first family of purely dense, decoder-only reasoning large language models, available in 3B, 8B, and 30B sizes. Each model was pretrained from scratch on approximately 15 trillion (15T) tokens and uses a five-stage training strategy to extend the context window to 512K tokens. The models were then supervised fine-tuned with chain-of-thought, reasoning, and agent trajectory data, followed by post-training through a multi-stage reinforcement learning (RL) pipeline. This pipeline includes agentic RL, through which the 8B and 30B models learn to use tools and perform actions in real sandbox environments. Every model supports switching between thinking and non-thinking modes, as well as a low-effort thinking mode that uses a shorter reasoning budget for simple questions. All models natively support tool calling. All Granite 4.2 models are released under the Apache 2.0 license.

Overview

Granite 4.2 is the reasoning-focused version of the Granite language model series. Earlier Granite models were capable instruction-following assistants; Granite 4.2 adds explicit reasoning capabilities. Each model can generate a chain of thought before answering and can run in either thinking or non-thinking mode, depending on the level of reasoning required by the task. The low-effort mode lies between the two, using a shorter reasoning budget for simple questions.

The three sizes (3B, 8B, and 30B) use the same architectural design and follow the same training pipeline—from-scratch pretraining, SFT, and then multi-stage RL—with model size as the only difference. All three models offer strong reasoning and instruction-following capabilities. The most significant capability differences emerge during post-training. The 8B and 30B models also undergo an additional agentic reinforcement learning (agentic RL) stage, where they learn to operate as agents: calling tools in real environments, editing and running code, operating terminals, and searching the web. All models support native tool calling. When served through an OpenAI-compatible endpoint, such as with vLLM, the models output tool calls in the OpenAI function-calling format, allowing them to integrate with agent frameworks without additional adaptation. Granite 4.2 also supports SGLang. See the SGLang cookbook for deployment-ready configurations.

The remainder of this article describes how the models were built, including their architecture, pretraining, supervised fine-tuning, multi-stage RL pipeline, and evaluation results.

Model Architecture

Granite 4.2 models are based on a purely decoder-only, dense Transformer architecture with the following core components:

  • Attention mechanism: Grouped Query Attention (GQA), with 40 attention heads and 8 KV heads
  • Position embeddings: Rotary Position Embedding (RoPE), θ = 10,000,000
  • Feed-forward network: MLP with SwiGLU activation
  • Normalization: RMSNorm (ε = 1e-5)
  • Embedding layers: Input and output embeddings are separate and do not share weights
  • Precision: bfloat16
Component3B Dense8B Dense30B Dense
Embedding dimension256040964096
Number of layers404064
Attention head dimension64128128
Number of attention heads403232
Number of KV heads888
MLP hidden dimension81921280032768
MLP activation functionSwiGLUSwiGLUSwiGLU
Sequence length131072131072131072
Position embeddingsRoPERoPERoPE
Number of parameters3B8B30B

Pretraining

Granite 4.2 was trained from scratch on approximately 15 trillion (15T) tokens using a five-stage training strategy. Stages 1–2 focused on foundational pretraining. Stages 3–4 consisted of mid-training, with progressively higher-quality data introduced through annealing. Stage 5 introduced long-context training, extending the context window to 512K tokens. Each stage used a different data mixture and learning-rate schedule, gradually shifting the training data from broad, web-scale sources toward more rigorously curated, high-quality data sources.

The overall pretraining approach continues that of the previous-generation models. For a detailed discussion of the data mixture, stage organization, and long-context extension, see the Granite 4.1 blog.

SFT: Data Preparation and Quality Control

Supervised fine-tuning (SFT) transforms the base model into a reliable instruction-following, reasoning, and tool-use assistant. The SFT data mixture contains agentic data (31.6%) and non-agentic data (68.4%), totaling approximately 7.2 million samples, or about 100B tokens, of which approximately 65B tokens are used for training.

The agentic corpus covers multiple domains, including software engineering (SWE, 69%), tool calling (12.1%), terminal use (8.0%), mathematics (3.5%), search (0.8%), and action execution (0.2%). These samples and trajectories were generated by a variety of agent scaffolds and execution frameworks, including OpenHands, OpenCode, Terminus-2, SWE-agent, OpenResearcher, MiniSWE, OpenSeeker, EnvScaler, Gemini CLI, Hermes, Codex, and Goose. The agentic data includes samples from open-source datasets as well as data generated in our proprietary synthetic RL environments, covering various combinations of agents and execution frameworks.

The non-agentic corpus consists of several major categories: instruction following (18.8%), code (18.8%), mathematics (14.6%), multilingual data (7.0%), science (5.4%), reasoning (3.0%), and safety (0.8%).

Data Quality Control

We perform multiple stages of quality control before samples enter the final SFT data mixture. First, data from different sources is standardized and reformatted into a consistent OpenAI Chat format, ensuring that conversation structures and tool interactions remain consistent across datasets and scaffolds.

We then use GPT-OSS-120B and Gemma 4 as large-language-model-based judges to evaluate sample quality. Low-scoring samples are removed. Samples containing hallucinations or fabricated information, invalid tool interactions, or calls to functions not defined in the corresponding tool list are also removed. Where applicable, we apply additional dataset-specific heuristics to further improve quality and eliminate known sources of noise.

Finally, we perform both local and global deduplication. Deduplication is based on computing a SHA-256 hash over the combination of the tools and messages fields, removing duplicate samples both within individual data sources and across the entire SFT data mixture.

SFT Training Details

The complete corpus is first globally shuffled to reduce the impact of data ordering and ensure that samples from different domains are thoroughly mixed throughout training. The shuffled corpus is then divided into equally sized .parquet shards, tokenized using the model’s tokenizer and chat template, and prepared for large-scale distributed training.

Before launching the final large-scale training run, we tune hyperparameters on representative configurations, scanning learning-rate schedules, initial learning rates, and warmup ratios to identify settings that train stably across different model sizes. The final training configuration is as follows:

ParameterValue
Compute resources32–128 nodes, depending on model size; each node is equipped with 4 Grace/GB200
Sequence length (after packing)131,072 (128K)
Global batch size128
Learning rate1.0e-5, held constant after warmup; 3.0e-6 for Stage 2
Learning-rate warmup2.5% of train_iters
Training durationApproximately 2 epochs
ParallelismTP=2, PP=1, CP=4 or CP=2

Stage 2 SFT for the 30B Model

For the 30B model, we also perform an additional stage