VibeThinker-3B: Exploring the Frontier of Verifiable Reasoning in Small Language Models | Summary
15 Jun 2026 | Paper Review Reasoning Reinforcement Learning Self-Distillation Test-Time ScalingContents
This article explains the key points of VibeThinker-3B: Exploring the Frontier of Verifiable Reasoning in Small Language Models.
- 2026-06-15 (arXiv)
- Xu, Sen, Liu, Shixi, Wang, Wei, Min, Jixin, Dai, Yingwei, Yin, Zhibin, Chen, Yirong, Zhou, Xin, Zhang, Junlin.
- Sina Weibo Inc.
- Paper
- Github
- Project Page
Summary
- VibeThinker-3B: Exploring the Frontier of Verifiable Reasoning in Small Language Models investigates whether a dense 3B-parameter model can approach leading reasoning systems on mathematics and executable coding. Built on Qwen2.5-Coder-3B base, it combines curriculum-based supervised fine-tuning, multi-domain reinforcement learning, offline self-distillation, and instruction-oriented reinforcement learning.
- The pipeline follows the Spectrum-to-Signal Principle: supervised learning preserves diverse valid solution paths, while MaxEnt-Guided Policy Optimization emphasizes verifiable problems near the model’s current capability boundary. Reasoning RL uses a single 64K context window, and Long2Short Math RL redistributes rewards among correct trajectories to favor concise reasoning.
- VibeThinker-3B reports 94.3 on AIME26, 80.2 Pass@1 on LiveCodeBench v6, and 93.4 on IFEval. Claim-Level Reliability Assessment (CLR), an inference-time sampling and self-verification procedure, raises AIME26 to 97.1 and IMO-AnswerBench from 76.4 to 80.6; recent LeetCode contests yield 96.1% acceptance across 128 first-attempt Python submissions.
The six panels show that competitiveness varies by benchmark. CLR raises AIME26 from 94.3 to 97.1, while baseline LiveCodeBench v6 at 80.2 remains below every displayed flagship comparator. IFBench at 74.5 exceeds four of the five displayed comparators but trails GLM-5 at 76.5, supporting task-specific strength rather than uniform superiority.
- These results show strong performance per parameter on structured, verifiable tasks, not equivalence to leading general-purpose models. GPQA-Diamond remains at 70.2 without CLR and 72.9 with CLR, and the report provides neither component ablations nor compute-matched comparisons that establish which pipeline changes caused the gains.
1 Introduction
The research question extends VibeThinker-1.5B from demonstrating feasible small-model reasoning to examining the performance boundary at a strict 3B parameter budget. The authors focus on difficult mathematical derivations and programming tasks with reliable outcome verification.
- The 3B iteration changes both parameter scale and post-training: stricter data filtering, a two-stage SFT curriculum, multi-domain RL, long-context training, efficiency-oriented Math RL, self-distillation, and instruction alignment.
- The comparison with the earlier 1.5B model therefore does not isolate the effect of doubling parameter count.
The Parametric Compression-Coverage Hypothesis distinguishes parameter-dense capabilities, such as reusable search and constraint-solving procedures, from parameter-expansive capabilities requiring broad factual and semantic coverage. The related Reasoning-Knowledge Decoupling Paradigm positions compact reasoning models as complementary to large generalists; Figure 2 presents the motivating parameter-performance comparison, but does not establish how these capabilities are encoded.
- On the 400-problem IMO-AnswerBench, VibeThinker-3B scores 76.4, or 80.6 with CLR, versus DeepSeek V3.2 at 78.3 with 671B parameters, GLM-5 at 82.5 with 744B, and Kimi K2.5 at 81.8 with 1T.
- The GPQA-Diamond gap is consistent with the hypothesis, but cross-model benchmark comparisons do not separate knowledge coverage from reasoning capacity.
VibeThinker-3B reaches 76.4 on the 400-problem IMO-AnswerBench, and CLR raises it to 80.6, above DeepSeek V3.2 at 78.3 but below GLM-5 at 82.5 and Kimi K2.5 at 81.8. The disclosed parameter counts demonstrate compactness, but the CLR extension requires additional inference work that the chart does not quantify.
2 Methods
Overall Pipeline. Figure 3 organizes post-training around Qwen2.5-Coder-3B base. Diversity-Exploring Distillation builds a broad solution spectrum during two-stage SFT; sequential Math, Code, and STEM RL strengthens verifiable reasoning; offline self-distillation consolidates stage-specific capabilities; and Instruct RL supplies final instruction-oriented alignment.
- The mathematical RL stage moves from accuracy optimization to Long2Short efficiency optimization.
- Checkpoints from the domain-specific RL stages supply verified trajectories for the subsequent unified student.
The diagram separates broad-coverage SFT from hard-reasoning SFT, with Diversity-Exploring Distillation in both stages. Reasoning RL proceeds through Math, Code, and STEM, and its checkpoints feed offline self-distillation before constraint- and rubric-based Instruct RL. The Math RL block marks the transition from accuracy to efficiency.
2.1 Supervised Fine-tuning
2.1.1 Data Construction. The SFT mixture covers math, code, STEM reasoning, general chat, and instruction following. Data Synthesis and Query Expansion. Mathematical seed queries require credible answers or rationales, and programming seeds require reliable tests or executable evaluation rules; expansion varies concept composition, solution skeletons, constraints, and evaluation objectives.
- Strong teacher models independently sample responses to initially filtered synthetic queries, with majority voting used to generate pseudo-labels.
- Multi-path Reasoning Distillation. Multiple candidate traces retain complete intermediate steps and diverse valid methods rather than reducing each problem to one standard solution.
Multi-level Quality Control. N-gram filtering removes repetitive or degenerate generations and overlaps with evaluation sets; LLM-based query filtering rejects incomplete, unreasonable, or logically invalid problems. Trace correctness filtering combines answer verification, sandbox execution, and LLM majority voting, after which data are stratified by reasoning length and difficulty.
- The pipeline explicitly includes benchmark decontamination and supervision checks, but the report does not specify dataset sizes, teacher identities, or detailed contamination thresholds.
2.1.2 Training Process. Curriculum-based two-stage SFT strategy. Stage one trains on the entire quality-filtered dataset with sequence packing, global batch size 128, initial learning rate 5 × 10⁻⁵, cosine decay to 8 × 10⁻⁸, 5 epochs, and 5% linear warmup. Stage two continues for 2 epochs with the same hyperparameters, retaining traces of at least 5K tokens and problems whose error rate is at least 0.75 across 8 VibeThinker-1.5B rollouts.
- Diversity-Exploring Distillation. In both SFT stages, intermediate checkpoints are evaluated by domain-specific Pass@K, selecting specialists for their ability to produce valid solutions rather than lowest validation loss or highest Pass@1.
- The specialists are merged at the parameter level to produce a unified SFT model; the report does not specify the merging rule or the probing-set K.
2.2 Reinforcement Learning
2.2.1 Algorithm Backbone. MaxEnt-Guided Policy Optimization (MGPO) weights each prompt according to the empirical correctness of a group of G sampled responses: (p(q)=\frac{1}{G}\sum_{i=1}^{G}\mathbb{I}(r_i=1)). Its prompt weight is w(q) = exp(−γ DME(p(q) ∥ p₀)), with p₀ = 0.5 and γ > 0, favoring groups containing both correct and incorrect trajectories.
- The weight multiplies the group-relative advantage inside a GRPO-style clipped, token-level probability-ratio objective.
- The authors describe DME as deviation from the maximum-entropy point 0.5, but do not give its explicit formula in this report.
- All RL stages are on-policy to mitigate training–inference probability mismatch, which the authors observed could destabilize or collapse training.
2.2.2 Multi-domain Reasoning RL. Training Data. Each domain uses benchmark-decontaminated samples with reliable supervision, excluding prompts whose starting-checkpoint accuracy is exactly 0.0 or 1.0. Mathematical rewards primarily verify final answers, code rewards execute sandbox tests, and STEM rewards combine answer matching with option verification.
Single Long-context Learning. Unlike the earlier 1.5B system, VibeThinker-3B uses a single 64K context window throughout reasoning RL. The authors report that early high-truncation training weakened long-thinking behavior and that later context expansion did not fully recover it.
- They hypothesize that the stronger SFT initialization already contains useful long-horizon traces, so truncation disrupts valid behavior instead of mainly suppressing noisy traces; no quantitative context-schedule ablation is provided.
- Training Strategy. RL proceeds sequentially through Math, Code, and STEM, targeting symbolic derivation and search, executable logic and boundary cases, and multidisciplinary scientific reasoning, respectively.
- Each domain-stage checkpoint is retained for offline self-distillation rather than relying exclusively on the final sequential checkpoint.
Long2Short Math RL. After accuracy-oriented MGPO, a second mathematical RL stage favors shorter correct responses while leaving incorrect rewards unchanged. For correct trajectories C, with sᵢ = 1/Lᵢ and mean brevity s̄, the reward becomes r′ᵢ = rᵢ + λ(sᵢ − s̄)/maxⱼ∈C|sⱼ − s̄|, with λ = 0.2.
- If all correct trajectories have the same length, their rewards remain unchanged.
- The centered shifts sum to zero over C, preserving both the correct-subset mean reward and the whole-group mean reward while changing relative preferences among correct paths.
- The objective is to preserve validation performance while removing redundant reasoning; the report provides no measured token reduction or accuracy–length trade-off.
2.3 Offline Self-Distillation
Offline Self-Distillation transfers verified trajectories from the Math, Code, and STEM RL checkpoints into one student through SFT. Learning-potential Filtering. After verifier-based rejection sampling, traces are scored by the student’s length-normalized negative log-likelihood, SLP(q, y) = −(1/|y|) ∑ₜ₌₁^|y| log πθstu(yₜ | q, y<t), prioritizing correct traces that the student does not yet model well.
- Selection occurs within domain-specific length buckets rather than through a global ranking.
- Extremely short traces and extreme high-score outliers are removed; verified middle-to-high-score traces are then mixed across domains.
- The filtering addresses abnormal-token sensitivity, format errors, distributional shifts, and noisy samples, but bucket boundaries and selection proportions are not specified.
2.4 Instruct RL
Instruct RL trains the reasoning-enhanced checkpoint on format-sensitive prompts, long-context instructions, and general alignment examples. Explicit constraints receive rule-based rewards for format, ordering, item count, keywords, and completion; open-ended prompts receive rubric-based rewards for helpfulness, coherence, instruction adherence, and redundancy.
- Both reward sources operate within the same on-policy RL framework.
- The final IFEval and IFBench results show strong constraint-following performance, but no before-and-after comparison isolates Instruct RL’s contribution or establishes that earlier capabilities were unchanged.
3 Evaluation
Evaluation examines the final 3B model across mathematics, executable coding, scientific knowledge, and instruction following. Standard benchmark results are supplemented by recent-contest coding results and a separate CLR-enhanced inference setting.
3.1 Evaluation Setup
Benchmarks. Mathematics includes AIME25, AIME26, HMMT25, BruMO25, and IMO-AnswerBench; coding includes LiveCodeBench v6 and OJBench. GPQA-Diamond measures graduate-level scientific reasoning, while IFEval and IFBench assess explicit instruction constraints.
- Evaluation protocol. VibeThinker-3B runs with vLLM, temperature 1.0, top-p = 0.95, and top-k = −1 unless otherwise specified, without an additional output-length cap beyond the model’s maximum generation length.
- Mathematical scoring jointly uses math verify and LLM-as-judge; coding correctness is determined by executing the generated solution against tests.
- Reported mathematical scores average Pass@1 over 64 independent generations, except IMO-AnswerBench, which uses 16; knowledge benchmarks use 16 generations and coding benchmarks use 8. Comparison-model scores come from released reports, public leaderboards, or official evaluation records rather than a uniform rerun.
Test-time scaling with claim-level reliability assessment. CLR generates K = 32 candidate trajectories under the standard sampling settings and extracts M = 5 decision-relevant claims alongside each final answer. The model self-verifies these claims with binary verdicts vₖ,ₘ, assigns reliability (r_k=\left(\frac{1}{M}\sum_{m=1}^{M}v_{k,m}\right)^M), and selects the answer-equivalence cluster maximizing Score(G) = ∑{k | yₖ∈G} rₖ.
- The nonlinear reliability score penalizes trajectories with rejected intermediate claims rather than weighting all sampled answers equally.
- The entire procedure is independently repeated 8 times, and its averaged final-answer Pass@1 is reported as “+ CLR”; this measures a multi-sample selection procedure, not a single generated trajectory.
- The authors claim lower token consumption than full-trace self-verification, but provide no comparative token counts, latency measurements, or verification-accuracy analysis.
3.2 Evaluation Results
Overview. The results address whether the complete 3B system can approach leading reasoning models, rather than whether increasing parameter count alone produces that outcome. Core benchmark performance. Table 1 reports 91.4 on AIME25, 94.3 on AIME26, 89.3 on HMMT25, 93.8 on BruMO25, and 76.4 on IMO-AnswerBench, exceeding every listed small-model baseline across this mathematics suite, including the 14B models.
- LiveCodeBench v6 reaches 80.2, the highest reported value in Table 1, whereas OJBench at 38.6 remains below LongCat Flash at 40.7.
- IFEval at 93.4 and IFBench at 74.5 are the highest reported values in Table 1. These establish strong final-model instruction performance, not preservation relative to earlier training stages.
Table 1 also shows that the lead does not extend to every evaluated domain. GPQA-Diamond is 70.2, below Qwen3.5-4B at 76.2 and Phi4-Reasoning-Plus at 81.9.
- This pattern is consistent with the authors’ proposed distinction between compact reasoning procedures and broad knowledge coverage, but differences in training data, initialization, and evaluation conditions prevent a causal interpretation.
VibeThinker-3B leads the listed small reasoning models across the mathematics suite and has the highest reported LiveCodeBench v6, IFEval, and IFBench scores in this comparison. OJBench trails LongCat Flash, and GPQA-Diamond trails several models, including Qwen3.5-4B, showing that the lead is not universal.
Comparison with top-tier reasoning models. Table 2 places baseline AIME26 at 94.3 beside DeepSeek V3.2 at 94.2 and Kimi K2.5 at 93.3. CLR raises the mathematical scores to 96.7 on AIME25, 97.1 on AIME26, 95.4 on HMMT25, 99.2 on BruMO25, and 80.6 on IMO-AnswerBench.
- CLR-enhanced AIME26 and BruMO25 exceed all reported comparison values in Table 2, but IMO-AnswerBench remains below GLM-5 at 82.5 and Qwen3.6 Plus at 83.8.
- Baseline LiveCodeBench v6 at 80.2 is close to DeepSeek V3.2 at 80.8, but below Gemini 3 Pro at 87.4; OJBench at 38.6 is below Gemini 3 Pro at 58.8.
- GPQA-Diamond improves from 70.2 to 72.9 with CLR, versus Gemini 3 Pro at 91.9. The results support benchmark-specific competitiveness, not comprehensive parity with flagship systems.
The separate baseline and “+ CLR” rows show that the strongest mathematical comparisons require additional sampling and self-verification. CLR yields 97.1 on AIME26 and 99.2 on BruMO25, while GPQA-Diamond remains at 72.9 and no CLR-enhanced coding results are reported. Mathematical competitiveness therefore does not imply general flagship parity.
Generalization test on recent LeetCode contests. Table 3 covers Apr. 25–May 31, 2026 using Python-only one-shot generation: four problems per contest and four independent rollouts per problem produce 16 first-attempt submissions per contest. Across BW181, BW182, BW183, W499, W500, W502, W503, and W504, VibeThinker-3B passes 123/128 submissions, or 96.1%.
- Its failures occur in BW183 at 12/16 and W503 at 15/16; the other six contests are 16/16.
- The aggregate sits between Gemini 3 Flash at 124/128 and GPT-5.2 at 122/128, below GPT-5.3-Codex at 128/128 and Gemini 3.1 Pro at 127/128.
- W501 is omitted because public LeetCode LLM ranking data are unavailable. The evaluation covers 32 distinct problems, not 128 distinct tasks; although the authors describe the contests as unseen, no training-data cutoff independently establishes that status.
VibeThinker-3B passes 123 of 128 independent first-attempt Python submissions across eight contests, with four rollouts for each of 32 distinct problems. Its aggregate is one accepted submission below Gemini 3 Flash and one above GPT-5.2; BW183 accounts for four of its five failures. These narrow differences should be interpreted in light of the limited number of distinct tasks.
4 Conclusion
The conclusion presents VibeThinker-3B as evidence that specialized post-training can bring a compact model into the performance range of representative frontier systems on several complex, verifiable tasks. It interprets the mathematics, coding, and recent-contest results through the Parametric Compression-Coverage Hypothesis and positions compact reasoning models as complementary to large general-purpose models.
- The evaluations demonstrate strong results for this particular 3B system; they do not establish a minimum parameter threshold for comparable reasoning.
- The proposed explanation—that reusable reasoning cores compress more readily than broad knowledge coverage—remains a hypothesis about parameter representations.
Appendix
- The supplied PDF contains no appendix; pages 11–14 contain the conclusion and references. The report identifies arXiv:2606.16140v1, dated June 15, 2026, and provides project resources at https://github.com/WeiboAI/VibeThinker and https://huggingface.co/WeiboAI/VibeThinker-3B.
Brief Thoughts
The strongest evidence comes from results across different scoring mechanisms: AIME26 at 94.3, execution-verified LiveCodeBench v6 at 80.2, and constraint-based IFEval at 93.4. The weaker GPQA-Diamond and OJBench comparisons delimit the model’s strengths more clearly than the broad claim of frontier-level reasoning.
The main unresolved questions concern attribution and total efficiency. The report supplies SFT hyperparameters and reward formulations, but lacks component ablations, full training-data specifications, compute budgets, measured Long2Short savings, and compute-matched CLR baselines; performance per parameter should not be equated with demonstrated end-to-end efficiency. Comparison scores drawn from public sources and the absence of uncertainty intervals make small score differences insufficient to establish superiority. Stage-by-stage evaluation and a comparison of CLR with ordinary answer voting under equal token budgets would clarify the contributions of training and inference.