Gorio Tech Blog search

Ministral 3 | Summary

|

Contents

This article explains the key points of Ministral 3.

  • 2026-01-13 (arXiv)
  • Liu, Alexander H., Khandelwal, Kartik, Subramanian, Sandeep, Jouault, Victor, Rastogi, Abhinav, Sadé, Adrien, Jeffares, Alan, Jiang, Albert, Cahill, Alexandre, Gavaudan, Alexandre, et al.
  • Paper
  • Project Page

Read this article in Korean


Summary

  • Ministral 3 introduces nine open-weight dense models: Base, Instruct, and Reasoning variants at 3B, 8B, and 14B parameter scales. All support image understanding and are released under Apache 2.0; the introduction specifies context lengths up to 256k tokens, with 128k for Reasoning models.
  • Cascade Distillation progressively prunes a pretrained Mistral Small 3.1 parent into smaller models and restores capability through logit distillation. The reported child-model training horizons are between 1 and 3 trillion tokens, but this comparison does not account for the parent model’s original training cost.
  • The 14B models show the strongest comparative results: Base retains much of its 24B parent’s capability, and Reasoning exceeds Qwen 3 14B on all six reported reasoning benchmarks. Smaller-model results are mixed, and teacher ablations show that a more capable teacher does not necessarily produce a better student during pretraining.

1 Introduction

The paper addresses how to produce capable compact models for compute- and memory-constrained applications without independently pretraining each size from scratch. Ministral 3 reuses knowledge and weights from the 24B Mistral Small 3.1 model to derive progressively smaller descendants.

  • The introduction contrasts the child models’ 1–3 trillion-token training horizons with 36 trillion tokens for Qwen3 and 15 trillion tokens for Llama3. These are reported training-horizon comparisons, not matched end-to-end compute measurements.
  • Figure 1 distinguishes the sequential pretraining cascade from the separate instruction-following and reasoning post-training branches.
  • The release resources are https://mistral.ai/news/mistral-3 and https://huggingface.co/collections/mistralai/ministral-3.

The diagram separates two forms of reuse: short-context child checkpoints initialize the next smaller model, while Mistral Small 3.1 supplies teacher logits throughout pretraining. Instruct and Reasoning are separate branches from Base, rather than consecutive stages of one post-training pipeline.

Cascade pretraining and separate SFT–ODPO and SFT–GRPO–ODPO post-training branches for the Ministral 3 family.
Cascade pretraining and separate SFT–ODPO and SFT–GRPO–ODPO post-training branches for the Ministral 3 family.

2 Model Architecture

The models use decoder-only Transformers with Grouped Query Attention, 32 query heads, 8 key-value heads, RoPE, SwiGLU, RMSNorm, and a 131K-token vocabulary. Table 1 shows reductions in both depth and width; only the 3B model ties its input and output embeddings.

  • In 14B, 8B, and 3B order, the layer counts are 40, 34, and 26; latent dimensions are 5120, 4096, and 3072; and FFN dimensions are 16384, 14336, and 9216.
  • Long-context extension uses YaRN and position-based softmax temperature scaling. Table 1 lists 256k contexts for the architectures, while the introduction separately limits the released Reasoning variants to 128k.
  • Every model inherits a frozen 410M-parameter ViT from Mistral Small 3.1 Base, using the architecture described in Pixtral. The original vision-to-language projection is discarded, and a new projection is trained for each model.

The specifications show reduced depth, latent width, and FFN width while preserving 32 query heads and 8 key-value heads. Tied embeddings distinguish the 3B model, limiting the parameters allocated to its 131K-token vocabulary.

Layer counts, latent dimensions, attention heads, FFN dimensions, embedding tying, and context lengths for the three model sizes.
Layer counts, latent dimensions, attention heads, FFN dimensions, embedding tying, and context lengths for the three model sizes.

3 Training Recipe

Training first produces multimodal Base checkpoints through pruning and continued pretraining with distillation. Each Base checkpoint then starts two separate post-training pipelines: SFT followed by ODPO for Instruct, and reasoning SFT followed by GRPO and ODPO for Reasoning.

3.1 Pretraining

Cascade Distillation follows a “prune-distill-repeat” sequence from Mistral Small 3.1 to 14B, then 8B, then 3B. Algorithm 1 preserves the short-context checkpoint as the starting point for the next pruning step, while a separate long-context training branch produces the final Base model at each size.

  • Mistral Small 3.1 remains the logit teacher throughout the cascade; the progressively smaller checkpoints supply initialization weights rather than replacing the original teacher.
  • The authors describe the process as continual pretraining with pruning en route, traversing the data mixture without repetition across the cascade.
  • Figure 2 illustrates sharp increases in cross-entropy loss after pruning, followed by recovery during continued training. It illustrates the cascade rather than quantifying an advantage over training from scratch.

Each pruning event is followed by a sharp loss increase and subsequent recovery during distilled training. The smaller checkpoints settle above the teacher’s loss, illustrating incomplete recovery to teacher-level cross-entropy; the plot does not quantify savings relative to independent pretraining.

Illustration of cross-entropy loss recovery across the 14B, 8B, and 3B short-context cascade.
Illustration of cross-entropy loss recovery across the 14B, 8B, and 3B short-context cascade.

Algorithm 2 combines layer pruning, hidden-dimension pruning, and feedforward-dimension pruning using activations from a large calibration batch. These operations adapt a larger checkpoint to the target architecture while retaining components selected by activation-based criteria.

  • Layer selection uses activation-norm ratios as an importance proxy. Although the surrounding prose describes an input-to-output ratio, Algorithm 2 explicitly computes (output_norm / input_norm).mean() and retains the top-scoring layers.
  • Hidden-dimension pruning applies PCA to concatenated attention-normalization and feedforward-normalization inputs across layers, producing a shared network-wide rotation before dimensionality reduction.
  • Feedforward pruning scores the mean absolute gated activation silu(layer.ffn.w1.output_x) * layer.ffn.w3.output_x and retains the highest-scoring dimensions.

After pruning, each child is trained on text-only data and interleaved text-image data using forward KL logit distillation. The authors report that this objective alone outperforms tuning a weighted combination of distillation and next-token prediction losses, but provide no numerical loss-objective ablation.

  • The short-context stage uses a context window of 16,384 tokens, and its output initializes the next child model.
  • The long-context stage extends the window from 16,384 to 262,144 tokens using YaRN and position-based temperature scaling.

3.2 Post-Training: Ministral Instruct

3.2.1 Supervised Fine-tuning uses curated text-only and multimodal instruction data, fp8 quantization, and a logit distillation loss from Mistral Medium 3. The authors report benefits from a stronger teacher at this stage, unlike pretraining; the vision encoder remains frozen while the adapter is trainable.

3.2.2 Online Direct Preference Optimization stage samples two responses from the current policy at T=0.7 and ranks them with a text-based Pairwise Reward Model (PWRM). Rather than using hard winner/loser labels, training uses a two-sided DPO loss weighted by the PWRM’s predicted preference probabilities.

  • The PWRM is trained by SFT on conversation histories with paired candidate responses. Its temperature is adjusted to calibrate preference probabilities, and β-rescaling is used to stabilize the DPO loss.
  • Responses exhibiting infinite loops are automatically treated as losers to discourage this generation artifact.
  • Tool execution is enabled during generation. The authors report improved tool-use performance and stronger preference alignment than SFT or offline preference optimization, without supplying numerical ablations for these claims.

3.3 Post-Training: Ministral Reasoning

3.3.1 Reasoning Supervised Fine-Tuning begins from the long-context pretrained checkpoint, not the Instruct ODPO checkpoint. It mixes short and long chain-of-thought samples across mathematics, coding, dialogue, instruction following, multilingual tasks, tool use, and visual reasoning.

  • Long reasoning traces are prefixed with a reasoning-specific system prompt. Filtering removes poor formatting, excessive repetition, and undesirable language switching.
  • For 3B, ordinary SFT produced brittle, repetitive, overly verbose outputs and infinite generations. The authors report that logit distillation from Magistral Small 1.2 reduced verbosity and stabilized subsequent RL.

3.3.2 Reinforcement Learning applies GRPO in two stages: STEM RL on math, code, and visual reasoning, followed by General RL on broader conversational and instruction-following prompts. General RL uses atomic grading rubrics, with an LLM judge assigning a reward equal to the fraction of satisfied heuristics.

  • STEM question-answer pairs come from open and proprietary sources and are filtered to remove invalid, incomplete, and very easy or very hard problems. Further recipe details are delegated to Rastogi et al. [2025].
  • The authors report that General RL improves chat and instruction following while maintaining, and sometimes improving, STEM performance, but do not provide a stage-by-stage numerical table.
  • Maximum generation length is increased from 32K to 80K because a non-trivial proportion of RL generations were truncated. The authors report additional performance gains when difficult reasoning trajectories can finish.

3.3.3 Online Direct Preference Optimization adds a final preference-alignment stage after GRPO. It follows the Instruct ODPO setup, except that thinking chunks are removed before the reward model scores each generated response.

  • This stage targets conversational and instructional behavior rather than directly scoring the removed reasoning trace.
  • Its size-dependent results are examined in Section 5.3 and Figure 6.

4 Results

The evaluation covers general knowledge and reasoning, mathematics and coding, multilingual knowledge, multimodal understanding, and post-training behavior. External baselines are rerun with the authors’ evaluation pipeline, and Table 2 explicitly states that its configurations are identical across models.

  • Base evaluations include MMLU, MMLU-Redux, ARC-Challenge, RACE High, TriviaQA, NaturalQS, AGIEval, MATH, GPQA Diamond, MBPP, MMMU, and MathVista.
  • Post-training evaluations include Arena Hard, WildBench, MM MTBench, AIME 2024/2025, HMMT 2025, PhyBench, and LiveCodeBench.
  • MM MTBench is linked at https://huggingface.co/datasets/mistralai/MM-MT-Bench. The report does not provide a complete reproducible evaluation configuration for every benchmark.

4.1 Pretraining Results

Table 2 shows competitive but benchmark-dependent Base results. Ministral 3 14B exceeds Qwen 3 14B on TriviaQA, 74.9 versus 70.3, and MATH, 67.6 versus 62.0, while scoring lower on MMLU-Redux, AGIEval, and Multilingual MMLU.

  • At 8B, Ministral scores 79.3 versus Qwen’s 79.4 on MMLU-Redux, 68.1 versus 63.9 on TriviaQA, and 62.6 versus 57.6 on MATH. It also exceeds Gemma 3 12B on four of the five listed metrics, but not TriviaQA.
  • At 3B, Ministral exceeds Qwen 3 4B on TriviaQA, 59.2 versus 53.0, and MATH, 60.1 versus 40.5, but trails on the other three listed metrics.
  • The text claims that Ministral 3 14B is better than Gemma 3 12B across all benchmarks, but Table 2 contradicts that statement on TriviaQA: 74.9 versus 78.8.

The results show higher MATH and TriviaQA scores than Qwen at each listed scale, but not uniform superiority across benchmarks. Gemma 3 12B scores 78.8 on TriviaQA versus Ministral 3 14B’s 74.9, contradicting the surrounding all-benchmarks claim.

Base-model comparisons on MMLU-Redux, TriviaQA, MATH, AGIEval, and Multilingual MMLU under identical internal-harness configurations.
Base-model comparisons on MMLU-Redux, TriviaQA, MATH, AGIEval, and Multilingual MMLU under identical internal-harness configurations.

Table 3 compares the children directly with their 24B parent and shows substantial capability retention alongside losses on some tasks. Ministral 3 14B nearly matches the parent on MMLU-Redux, 82.0 versus 82.7, and matches MBPP at 71.6, while exceeding it on MATH, 67.6 versus 55.8.

  • All three children exceed the parent’s MATH score: 67.6, 62.6, and 60.1 for 14B, 8B, and 3B, respectively. This does not imply uniformly better reasoning, since GPQA Diamond falls to 33.8 at 3B versus the parent’s 36.9.
  • Knowledge retrieval degrades with shrinking size: NaturalQS falls from 34.4 for the parent to 29.9, 25.8, and 21.9.
  • The multimodal results are uneven. The 14B model scores 59.9 versus the parent’s 59.1 on MMMU, but MathVista falls from 51.3 to 43.6, 35.7, and 23.3 across the three child sizes.

The children retain much of the parent’s performance and all exceed its MATH score, but shrinking size produces clear declines in NaturalQS and MathVista. The 14B model exceeds the parent on MMMU, 59.9 versus 59.1, while trailing it on MathVista, 43.6 versus 51.3, showing task-dependent multimodal retention.

Mistral Small 3.1 24B and Ministral 3 Base results across general, math and code, multilingual, and multimodal benchmarks.
Mistral Small 3.1 24B and Ministral 3 Base results across general, math and code, multilingual, and multimodal benchmarks.

4.2 Post-training Results

Table 4 reports strong 14B Instruct results and mixed comparisons at smaller sizes. Ministral 3 14B exceeds Qwen3 14B (Non-Thinking) on Arena Hard, 55.1 versus 42.7; WildBench, 68.5 versus 65.1; and MATH (maj@1), 90.40 versus 87.00.

  • Against Qwen3-VL-8B-Instruct, Ministral 3 8B is slightly higher on WildBench and MM MTBench, but lower on Arena Hard, 50.9 versus 52.8, and MATH, 87.60 versus 94.60.
  • Ministral 3 3B exceeds Qwen3-VL-2B-Instruct on all four listed metrics, but trails Qwen3-VL-4B-Instruct on three and ties WildBench at 56.8.
  • Although the surrounding text emphasizes vision-enabled Qwen3-VL baselines, the 14B row is Qwen3 14B (Non-Thinking) and has no MM MTBench result.

The 14B Instruct model exceeds its listed Qwen3 non-thinking baseline on all available shared metrics. At 8B, the comparison with Qwen3-VL is mixed, while 3B exceeds the listed 2B baseline but trails the 4B baseline on three metrics and ties it on WildBench.

Instruct-model comparisons on Arena Hard, WildBench, MATH (maj@1), and MM MTBench.
Instruct-model comparisons on Arena Hard, WildBench, MATH (maj@1), and MM MTBench.

Table 5 shows that Ministral 3 14B Reasoning exceeds Qwen 3 14B on all six reported benchmarks, including AIME 2025 at 85.0 versus 73.7 and LiveCodeBench v6 at 64.6 versus 59.3. The report states that reasoning evaluations use pass@16, except LiveCodeBench, which uses pass@5.

  • At 8B, Ministral ties Qwen3-VL on AIME 2024 at 86.0 and exceeds it on LiveCodeBench v6, 61.6 versus 58.0, but trails on the other four benchmarks.
  • Ministral 3 3B exceeds Qwen3-VL 4B on five of six benchmarks, including AIME 2025 at 72.1 versus 69.7, but is substantially lower on GPQA Diamond, 53.4 versus 60.1.
  • These comparisons use the same evaluation pipeline, but the 3B–4B comparison is not exactly parameter-matched. The stated multi-sample evaluation should not be presented as single-generation accuracy.

The 14B Reasoning model leads its Qwen 3 counterpart on all six metrics. The 8B model ties AIME 2024, leads on coding, and trails on the remaining four benchmarks; 3B exceeds the 4B baseline on five metrics while trailing on GPQA Diamond. The paper states pass@16 evaluation except pass@5 for LiveCodeBench.

Reasoning-model results on AIME 2024, AIME 2025, HMMT 2025, GPQA Diamond, PhyBench, and LiveCodeBench v6.
Reasoning-model results on AIME 2024, AIME 2025, HMMT 2025, GPQA Diamond, PhyBench, and LiveCodeBench v6.

5 Discussions

The discussion examines teacher selection, response verbosity, and preference alignment after reasoning RL. It combines plotted ablations with qualitative training observations; several reported improvements lack numerical results or detailed experimental settings.

5.1 Choice of Teacher Model for Distillation

Figure 3 shows that the 14B student performs better when pretrained through distillation from Mistral Small 3.1 than from the larger, stronger Mistral Medium 3 teacher. The authors relate this result to a teacher–student capacity gap, noting that the comparison is not FLOP-matched.

  • The plotted comparisons cover MMLU REDUX, HellaSwag, TriviaQA, MATH, AGI-EVAL, and Winogrande; the Small 3.1 teacher produces the higher-scoring student in each case.
  • Post-training behaves differently: the discussion reports benefits from the more capable Mistral Medium 3.1. Section 3.2 names Mistral Medium 3 for Instruct SFT, so the report does not consistently identify the exact Medium checkpoint.
  • The result supports testing teacher choice separately at each training stage; it does not establish that larger teachers generally underperform during pretraining.

Mistral Small 3.1 produces a higher-scoring 14B student than Mistral Medium 3 on every plotted benchmark. This supports the observed teacher-selection result, although the non-FLOP-matched setup does not isolate its mechanism or establish a general rule about teacher efficiency.

14B pretraining teacher ablation comparing Mistral Small 3.1 and Mistral Medium 3 across six benchmarks.
14B pretraining teacher ablation comparing Mistral Small 3.1 and Mistral Medium 3 across six benchmarks.

Figure 4 shows that distilling a 3B student from a post-trained Mistral Small 3.1 teacher improves MATH and MBPP relative to a Base teacher, with much smaller changes on knowledge and multimodal metrics. The plotted legend specifically compares MS3.1 Base with MS3.1 Instruct 2506.

  • The largest visible difference is on MATH (maj@4); MMLU REDUX, HellaSwag, and TriviaQA remain close, while MMMU improves modestly.
  • The authors also compare internal Mistral Medium 3 SFT and preference-tuned teachers during student SFT. They report that the preference-tuned teacher is consistently better and that its gains persist after the student’s own preference tuning.
  • No corresponding numerical table, uncertainty estimate, or detailed setup is provided for the preference-tuned-teacher comparison.

The Instruct 2506 teacher’s clearest advantage appears on MATH (maj@4), with additional gains on MBPP and a modest MMMU increase. Nearly unchanged knowledge scores indicate that the benefits are concentrated rather than uniform across capabilities.

3B pretraining ablation comparing MS3.1 Base and MS3.1 Instruct 2506 teachers.
3B pretraining ablation comparing MS3.1 Base and MS3.1 Instruct 2506 teachers.

5.2 Model Verbosity.

Figure 5 plots GPQA Diamond accuracy against output-token count and shows Ministral Instruct models producing substantially shorter responses than the plotted Qwen3-VL Instruct baselines. The authors suggest that differences in reasoning RL contribute to this behavior, but do not test that explanation through a controlled ablation.

  • Increasing the proportion of long CoT traces in Instruct SFT reportedly improves STEM benchmark performance but also introduces excessive reflection, internal monologues, and backtracking.
  • The boxed example illustrates repeated reconsideration within a mathematical answer. It supports the qualitative concern about chat style, not a measured frequency of the behavior.
  • Despite the caption’s reference to instruction-following and reasoning, the visible point labels in Figure 5 identify Instruct models.

The Ministral Instruct points appear far to the left of the Qwen3-VL Instruct points, indicating shorter GPQA Diamond responses. Ministral 14B combines the highest plotted accuracy with a comparatively low output-token count, but this task-specific plot does not directly measure latency or energy use.

GPQA Diamond accuracy versus number of output tokens for the plotted Instruct models.
GPQA Diamond accuracy versus number of output tokens for the plotted Instruct models.

5.3 ODPO for Ministral 3 Reasoning.

Figure 6 shows improvements on ArenaHard, EQBench, and WildBench when ODPO is applied to the 14B and 8B GRPO-trained Reasoning checkpoints. The 3B model shows much smaller public-benchmark changes, although internal human evaluations favored its ODPO checkpoint, which was selected for release.

  • The reward model scores responses after thinking chunks are removed, directing preference feedback toward the user-facing answer.
  • The authors report that the 3B Base model is more sensitive to fine-tuning hyperparameters than the 14B and 8B models.
  • The report does not describe the internal human-evaluation protocol or quantify the stated improvement, limiting independent assessment of the 3B release decision.

6 Conclusion

The paper concludes that iterative pruning and distillation can produce a compact family of multimodal Base, Instruct, and Reasoning models from existing larger checkpoints. The release provides Apache 2.0 open weights, but the conclusion’s general statement about 256K contexts should be read together with the introduction’s explicit 128k limit for Reasoning variants.

Appendix

  • The supplied PDF contains no appendix. After the conclusion and contributor lists, pages 12–14 contain references; no additional implementation details or experimental tables appear there.

Brief Thoughts

A useful feature of the recipe is its separation of initialization inheritance from teacher selection: smaller models inherit weights sequentially, while the original parent remains the pretraining teacher. The teacher ablations and the distinct behavior of 3B during reasoning SFT and ODPO make a case for evaluating compression and post-training choices at each model size.

The benchmark evidence supports competitive compact models, especially at 14B, more clearly than it establishes a precise compute-efficiency advantage. The report omits measured training FLOPs, wall-clock or memory comparisons, exact per-model token budgets, complete data mixtures, and confidence intervals; it also does not evaluate long-context quality at the advertised limits. The 1–3 trillion-token horizons benefit from an already pretrained parent, so they are not a complete accounting of the compute needed to create the family. Table 2’s TriviaQA exception and the mixed 8B post-training results warrant narrower claims than uniform superiority over similarly sized baselines. The increase to an 80K generation limit during reasoning RL also distinguishes parameter efficiency from inference-token efficiency; the limit is not a measurement of typical response length.