Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation | Summary
12 Feb 2026 | Paper Review On-Policy Distillation Reward Extrapolation Reinforcement LearningContents
- Summary
- 1 Introduction
- 2 Related Work
- 3 Methodology
- 4.1 Experiments with Same-Sized Student and Teacher
- 4.2 Experiments in the Strong-to-Weak Distillation Setting
- 5 Conclusion and Discussion
- Appendix
- Brief Thoughts
This article explains the key points of Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation.
- 2026-02-12 (arXiv)
- Yang, Wenkai, Liu, Weijie, Xie, Ruobing, Yang, Kai, Yang, Saiyong, Lin, Yankai.
- Gaoling School of Artificial Intelligence, Renmin University of China, LLM Department, Tencent
- Paper
- Github
Summary
- The paper reformulates on-policy distillation (OPD) as KL-constrained reinforcement learning with teacher-to-reference log-probability rewards that decompose across tokens. Generalized On-Policy Distillation (G-OPD) introduces a reward scaling factor λ and a selectable reference model; standard OPD is recovered at λ = 1.
- Reward extrapolation, termed ExOPD, uses λ > 1 to strengthen the implicit reward relative to KL regularization. With λ = 1.25, a unified Qwen3-4B-Non-Thinking student exceeds the corresponding math and code domain teachers on all seven benchmarks in the primary multi-teacher experiment, reaching average accuracies of 47.7% and 62.0%, respectively.
- ExOPD also improves math distillation from Qwen3-30B-A3B-Instruct-2507 into 1.7B and 4B students, while using the teacher’s pre-RL model as the reference provides additional gains in separate 4B-to-1.7B experiments. Extrapolation requires reference-model computation and generally produces longer responses; excessive extrapolation can degrade accuracy, and cross-family transfer remains untested.
The multi-teacher panel places ExOPD above both domain-teacher averages, whereas ExPO combines lower math accuracy with higher code accuracy. The strong-to-weak panel shows gains over OPD for both student sizes, not evidence that either student surpasses the larger teacher.
1 Introduction
OPD supervises student-generated trajectories with teacher predictions, addressing the mismatch between teacher-generated training data and the student’s inference behavior. The paper examines OPD’s connection to RL, whether changing its objective can improve on teacher imitation, and how reference selection affects transfer between models of different capacities.
- The first setting merges domain-specific RL variants of a common base model back into that base model.
- The second setting transfers capabilities from a larger teacher into a smaller student.
- The released implementation is available at https://github.com/RUCBM/G-OPD.
2 Related Work
The related work distinguishes distillation methods by the policy generating their training trajectories. G-OPD builds on teacher supervision of student-generated trajectories rather than proposing a new teacher-generation or black-box distillation mechanism.
Off-Policy Distillation.
Off-policy distillation trains on teacher-generated trajectories, either by matching teacher logits through KL divergence or by applying supervised fine-tuning (SFT) to sampled teacher tokens. SFT avoids requiring the teacher’s full output distribution but does not directly supervise trajectories sampled from the student’s current policy.
On-Policy Distillation.
OPD samples trajectories from the student and aligns its token distributions with the teacher, providing dense on-policy supervision. The cited literature also explores cross-family transfer, black-box teacher access, and self-distillation of contextual information; these are related directions, not settings evaluated here.
3 Methodology
The method starts from an objective-level equivalence between reverse-KL distillation and reward maximization with KL regularization. It then varies the reward-to-regularization weighting and the reference distribution used to define the implicit reward.
3.1 Preliminaries
The preliminaries contrast teacher-sampled knowledge distillation, student-sampled KL-constrained RL, and student-sampled reverse-KL OPD. The knowledge-distillation formulation matches the teacher to the student through forward KL, whereas OPD minimizes reverse KL from the student to the teacher.
- The RL objective is the expected response reward minus β times the KL divergence from a reference policy, which limits distributional drift.
- The reward may come from a learned preference model or a deterministic outcome verifier; the paper contrasts commonly used terminal rewards with OPD’s token-level supervision.
- Teacher-generated SFT is the practical off-policy alternative when the teacher’s full output distribution is unavailable or expensive.
The exact OPD gradient includes the cumulative student–teacher log-probability difference from the current token through the end of the response. The practical approximation uses a discount factor of 0 and retains only the current token’s difference.
- The local update is therefore an approximation to the sequence-level reverse-KL gradient, not the exact gradient of that objective.
- Appendix A shows why earlier-token terms vanish in expectation and derives the remaining suffix sum.
3.2 Generalized On-Policy Distillation
Introducing a reference policy decomposes standard OPD into the implicit reward log[π*(y|x)/πref(y|x)] and a KL penalty against πref with equal weights. G-OPD changes this balance through the objective JG-OPD(θ) = maxθ E[λ log(π*(y|x)/πref(y|x)) − DKL(πθ ∥ πref)].
- At λ = 1, the reference-dependent terms cancel, recovering standard OPD regardless of the reference choice.
- For λ ≠ 1, the reference affects the objective and its log-probabilities must be computed.
- By default, πref is the student’s initial policy; for positive λ, reward rescaling makes λ correspond to the inverse KL coefficient.
Equation (12) expresses the proposed target shift as log πθ(y|x) = λ log π*(y|x) + (1 − λ) log πref(y|x). The paper interprets 0 < λ < 1 as reward interpolation and λ > 1 as reward extrapolation, adding a teacher-to-reference shift beyond teacher matching.
- The authors conjecture that interpolation produces behavior between the reference and standard OPD, including intermediate response lengths.
- Extrapolation strengthens a proxy reward and can amplify its biases; it does not guarantee improved task accuracy.
- Equation (12), as printed, omits an input-dependent normalization term and should be read as an unnormalized log-policy relation.
For domain experts derived from a shared base, choosing that base as πref isolates the distributional change induced by domain-specific RL. In strong-to-weak transfer, reward correction replaces the student-base reference with the teacher’s pre-RL base, separating the teacher’s post-training shift from differences between teacher and student base models.
- Equation (13) rewrites G-OPD as teacher-directed reverse-KL regularization plus an additional implicit reward weighted by λ − 1.
- Reward correction requires the teacher’s pre-RL checkpoint and increases computation when this reference is larger than the student.
- The token coefficient printed in Equation (14) has the sign of a negative-objective loss gradient, rather than an ascent advantage for the maximization objective in Equation (11).
4.1 Experiments with Same-Sized Student and Teacher
The experiments first vary λ in same-sized single-teacher distillation, then merge math and code teachers, and finally evaluate strong-to-weak transfer. The same-sized setting uses Qwen3-4B-Non-Thinking as both the initial student and the common ancestor of the domain teachers.
4.1.1 Experimental Settings
Math training uses 57K DeepMath samples with difficulty greater than or equal to 6; code training uses 25K Eurus-RL-Code samples. Separate GRPO runs produce Qwen3-4B-Non-Thinking-RL-Math and Qwen3-4B-Non-Thinking-RL-Code, with reward 1.0 for a correct math answer or passing all code unit tests and 0.0 otherwise.
- The primary math and code teachers receive 500 and 300 optimization steps, respectively; both use batch size 128, rollout n = 8, learning rate 1 × 10−6, and KL coefficient 0.0.
- Maximum RL response lengths are 16,384 for math and 8192 for code.
- Distillation uses the same datasets and tests λ ∈ {0.0, 0.25, 0.5, 0.75, 1.0, 1.25, 1.5}, with the original student fixed as the reference; λ = 0.0 denotes the initial model.
- G-OPD uses batch size 1024, rollout n = 1, learning rate 1 × 10−5, and maximum response length 16,384. GRPO and G-OPD use verl and token-level rollout correction.
The math teacher uses 500 GRPO steps with eight rollouts per prompt and maximum response length 16,384. Its KL coefficient is 0.0, distinguishing the actual teacher-training configuration from the general KL-constrained formulation used in the theoretical discussion.
G-OPD uses batch size 1024, one rollout per prompt, and learning rate 1 × 10−5. These differ from the GRPO settings, so fewer distillation steps do not by themselves establish proportionally lower computation.
Math evaluation covers AIME24, AIME25, HMMT25 (February), and HMMT25 (November); code evaluation covers HumanEval+, MBPP+, and LiveCodeBench v6, February 2025∼May 2025. Evaluations use temperature 1.0, top-p 1.0, and maximum generation length 16,384, sampling 32 solutions per math problem and 4 per code problem before reporting average accuracy.
- Math answer verification uses Math-Verify: https://github.com/huggingface/Math-Verify.
- The metric averages sampled-solution correctness; it is not best-of-32 or best-of-4 success.
4.1.2 Results of Single-Teacher Distillation
The λ sweep shows standard OPD approximately recovering teacher accuracy and response length, while interpolation moves behavior toward the base model. At λ = 1.25, single-teacher ExOPD exceeds the primary teachers on every benchmark: average math accuracy is 48.0% versus 46.0% for the teacher and 46.5% for OPD, and average code accuracy is 62.1% versus 61.2% and 60.8%.
- Figures 2 and 3 show the benchmark-specific interpolation and extrapolation trends.
- The selected factor improves all seven primary results, but larger λ does not improve accuracy uniformly.
Math accuracy generally approaches teacher performance as λ increases toward 1. The λ = 1.25 setting improves all four primary math results, while further extrapolation is not uniformly beneficial.
The code results favor λ = 1.25 but show benchmark-dependent behavior at λ = 1.5. HumanEval+ declines at the larger factor, demonstrating that a higher reward weight does not necessarily improve accuracy.
Figure 4 shows that stronger reward weighting generally produces longer responses. At λ = 1.5, accuracy degrades on some benchmarks despite continued length growth; the authors suggest implicit-reward hacking and length bias, without directly testing these mechanisms.
- Longer outputs increase inference cost, and the experiments do not compare accuracy at a fixed generated-token budget.
Higher reward weighting generally shifts students toward longer outputs. Some λ = 1.5 points move farther right while losing accuracy, showing that extrapolation can increase generation cost without a corresponding performance gain.
Continued RL tests whether ExOPD’s advantage simply reflects insufficient teacher training: 100 extra math RL steps raise average accuracy from 46.0% to 46.9%, whereas 50 ExOPD steps reach 48.0%. Appendix C also evaluates teachers trained for 1200 RL steps; ExOPD retains higher domain averages, but does not exceed these stronger teachers on every math benchmark.
- With the 1200-step teachers, single-teacher ExOPD averages 52.2% in math and 64.3% in code, against teacher averages of 51.9% and 63.1%.
- Multi-teacher ExOPD averages 52.5% and 64.4%, while HMMT25 (November) remains below the math teacher at 42.7% versus 42.9%.
- Optimization-step counts do not establish equal training cost because RL and distillation use different rollout counts and model computation.
The teacher’s average math accuracy improves to 46.9% after 100 additional RL steps, while 50 ExOPD steps yield 48.0%. This weakens the insufficient-teacher-training explanation, although the different training procedures make step counts an incomplete cost comparison.
With 1200-step teachers, ExOPD maintains higher domain averages than OPD in both settings. Benchmark-wise exceptions remain: multi-teacher ExOPD scores 42.7% on HMMT25 (November) against the teacher’s 42.9%, and scores below OPD on AIME25 and HMMT25 (February).
4.1.3 Results of Multi-Teacher Distillation
The multi-teacher experiment fixes λ = 1.25 without further setting-specific tuning and downsamples math data to match the code sample count in OPD and ExOPD. SFT uses teacher-generated trajectories with matched trajectory counts, while ExPO averages domain-teacher weights and extrapolates them against the student using α selected from {0.25, 0.5}.
- OPD and ExOPD receive supervision from the teacher corresponding to each training domain.
- The same-sized G-OPD runs use 50 optimization steps.
- SFT matches the corresponding distillation step count and uses batch size 1024, maximum sequence length 32,768, warm-up ratio 0.05, and learning rate 1 × 10−5.
SFT uses the same batch size and learning rate as G-OPD, with maximum sequence length 32,768 and warm-up ratio 0.05. The accompanying text also matches trajectory counts and optimization steps, but does not equate the methods’ full computational costs.
In Table 2, multi-teacher ExOPD reaches 47.7% average math accuracy and 62.0% average code accuracy, exceeding OPD’s 46.4% and 60.6% and the domain teachers’ 46.0% and 61.2%. It is the only tested method in this primary experiment to exceed the corresponding teacher on every benchmark, although ExPO achieves a higher code average of 62.6%.
- SFT averages 44.3% in math and 60.8% in code, below both domain-teacher averages.
- ExPO averages 45.0% in math, so its stronger code average does not imply uniformly stronger transfer.
ExOPD exceeds the corresponding primary domain teacher on every benchmark in both single-teacher and multi-teacher distillation. ExPO nevertheless has the highest multi-teacher code average, so consistent teacher surpassing is distinct from outperforming every baseline on every metric.
Figure 5 shows ExOPD obtaining somewhat higher training rewards, longer responses, and higher response entropy than OPD. The curves use Exponential Moving Average smoothing with coefficient 0.5; the authors’ explanation that longer responses increase diversity is an interpretation of correlated trends.
ExOPD and OPD converge to similar training-reward ranges, with ExOPD typically somewhat higher. The clearer separation in response length, alongside higher entropy, is consistent with the longer outputs observed at evaluation.
4.2 Experiments in the Strong-to-Weak Distillation Setting
Strong-to-weak distillation evaluates extrapolation when the teacher is larger than the student rather than an RL variant of the student itself. The default setting requires the teacher and student base, whereas reward correction additionally requires the teacher’s pre-RL checkpoint.
4.2.1 Experimental Settings
The main strong-to-weak experiment distills Qwen3-30B-A3B-Instruct-2507 into Qwen3-1.7B-Non-Thinking and Qwen3-4B-Non-Thinking on the same math training data and four math benchmarks. ExOPD uses λ = 1.25 and the student base as its reference, with 100 distillation steps; comparisons include standard OPD and SFT.
- The authors cannot access the 30B teacher’s pre-RL checkpoint, so reward correction is not evaluated for that teacher.
- The reward-correction experiment instead uses the authors’ 4B domain teachers, whose pre-RL base is accessible.
4.2.2 Results of Strong-to-Weak Distillation
For the 1.7B student, ExOPD raises average math accuracy from OPD’s 23.1% to 25.4%, versus 13.5% for SFT and 8.8% for the base model. For the 4B student, the corresponding averages are 45.3%, 42.6%, 35.1%, and 15.4%; ExOPD improves on OPD on all four benchmarks for both students.
- The average improvements over OPD are 2.3 and 2.7 percentage points for the 1.7B and 4B students, respectively.
- Both remain below the teacher’s 59.7% average, demonstrating improved transfer rather than surpassing the larger teacher.
- The paper reports point estimates without confidence intervals or repeated-training variability.
ExOPD improves each math benchmark for both students relative to OPD, increasing their averages by 2.3 and 2.7 percentage points. The teacher’s 59.7% average remains above both students, limiting the teacher-surpassing result to other experimental settings.
4.2.3 Reward Correction in Strong-to-Weak Distillation
Reward correction uses Qwen3-1.7B-Non-Thinking as the student, the 4B math or code RL variant as teacher, and Qwen3-4B-Non-Thinking as the corrected reference. Average math accuracy rises from 28.1% with default ExOPD to 28.7% with correction, and average code accuracy rises from 51.3% to 52.3%.
- The corresponding OPD averages are 27.5% for math and 50.5% for code; SFT reaches 22.7% and 47.0%.
- These comparisons support the proposed reference choice in the controlled 4B-to-1.7B setting, but do not directly measure reward-signal accuracy.
- Correction requires an additional checkpoint and log-probability computation under a larger reference than default ExOPD.
Replacing the student-base reference with the teacher-base checkpoint improves ExOPD by 0.6 percentage points in math and 1.0 percentage point in code. These comparisons use accessible 4B domain teachers, not the 30B teacher from the main strong-to-weak experiment.
5 Conclusion and Discussion
The conclusion identifies reward weighting and reference selection as controls that can improve OPD beyond teacher matching in the evaluated settings. The authors leave validation on larger-scale models, broader collections of domain teachers, and different model families to future work.
Appendix
- A Detailed Math Derivations derives the reverse-KL OPD gradient using the score-function identity. The direct expected score term is zero, and log-probability differences from positions before the current action vanish in expectation, leaving a suffix-summed gradient before the discount-0 approximation; the appendix also states the corresponding local G-OPD gradient.
- B Detailed Training Settings provides the GRPO, G-OPD, and SFT hyperparameters. G-OPD uses 50 steps for same-sized pairs and 100 for strong-to-weak transfer; the authors report that further distillation can degrade generalization through overfitting and that larger prompt batches yield smoother convergence at a fixed prompt-count × rollout-n product.
- C Results of Distillation from Domain Teachers with Sufficient RL Trainings repeats the same-sized experiments with teachers trained for 1200 RL steps. Table 8 shows higher domain averages with ExOPD than OPD in both distillation settings, alongside exceptions to benchmark-wise teacher surpassing and benchmark-wise superiority over OPD.
- D Prompt Templates supplies the Training and Evaluation Prompt Template for Math Reasoning and the Training and Evaluation Prompt Template for Code Generation. The math template requests step-by-step reasoning and a final answer within \boxed{}; the code template requests reasoning followed by Python code in a final fenced block, with {question} as the input placeholder.
Brief Thoughts
The main empirical contribution is that λ = 1.25 improves standard OPD across the tested same-sized and strong-to-weak settings without further setting-specific tuning. The 1200-step teacher control strengthens this finding, but also shows why surpassing teachers on every benchmark is too broad a characterization.
The objective-level interpretation provides a useful design rationale, but the local gradient approximation and the normalization and sign issues in Equations (12) and (14) deserve care when reproducing the method. Longer responses, unreported statistical uncertainty, and unquantified reference-model overhead leave compute-matched effectiveness unresolved.