OPRD: On-Policy Representation Distillation | Summary
04 Jun 2026 | Paper Review On-Policy Distillation Representation Distillation Mathematical ReasoningContents
- Summary
- 1 Introduction
- 2 Background and Problem Setup
- 3 On-Policy Representation Distillation
- 4 Experiments
- 5 Discussion
- 6 Related Work
- 7 Conclusion and Future Work
- Appendix
- Brief Thoughts
This article explains the key points of OPRD: On-Policy Representation Distillation.
- 2026-06-04 (arXiv), Preprint
- Yang, Shenzhi, Zhu, Guangcheng, Song, Bowen, Wang, Haobo, Xia, Mingxuan, Zheng, Xing, Ma, Yingfan, Chen, Zhongqi, Wang, Weiqiang, Zhao, Junbo, et al.
- Zhejiang University, Ant Group
- Paper
- Github
Summary
- OPRD: On-Policy Representation Distillation replaces next-token distribution matching with intermediate hidden-state matching on student-generated responses. It uses the teacher’s layered representations directly rather than extracting supervision only from the LM-head output.
- With an aligned R1-distill-1.5B student and JustRL-1.5B teacher, OPRD-Vanilla reaches 49.8%, 34.6%, and 79.1% Avg@16 on AIME24, AIME25, and AIMO, leaving teacher gaps of 1.0, 1.0, and 0.4 percentage points. Over 500 optimizer steps on 8×A100 (80G), it takes 563 minutes rather than 812 minutes for OPD top-16 and reduces actor-update transient memory from 45.0 GB to 20.5 GB per GPU.
- OPRD-Bridge extends the method to heterogeneous models through frozen, layer-specific low-rank projectors, with reported transfer across architecture and tokenizer differences. In the cross-architecture experiment, it matches sampled-token OPD’s Best@16 but has lower Avg@16; its clearest advantages are shorter responses and less lexical repetition.
1 Introduction
The introduction motivates OPRD through token-sampling noise in sampled-token OPD and the lack of direct targets for intermediate representations in output-space objectives. OPRD retains student-generated trajectories but changes where teacher supervision is extracted.
- OPRD-Vanilla requires comparable representation coordinates, obtained experimentally from a teacher produced by RL training of the student’s original checkpoint. Matching architecture alone does not establish this comparability.
- OPRD-Bridge learns a shared representation interface when differences in depth, width, or training history make direct hidden-state MSE inappropriate.
- The implementation is available at https://github.com/ShenzhiYang2000/OPRD. Figure 1 summarizes the measured same-architecture accuracy, time, and transient-memory trade-off.
OPRD combines higher AIME24 accuracy with lower wall-clock time and a smaller transient-memory footprint than either evaluated OPD baseline. This is a measured Pareto improvement for the aligned 1.5B model pair and 500-step configuration, not a general result across architectures.
2 Background and Problem Setup
The background separates on-policy sampling from the choice of supervision space. It defines shared-vocabulary student and teacher models, compares sampled-token, full-vocabulary, and top-k output objectives, and identifies their common reliance on next-token distributions.
2.1 Notation
The student policy πθ has trainable parameters θ, while the teacher πT is fixed. Both process the same prompt and sampled response prefixes; intermediate hidden states are indexed by transformer layer and response position.
- L and d denote shared transformer depth and hidden dimension; LS, LT, dS, and dT distinguish heterogeneous models.
- The LM head maps the final hidden state into vocabulary logits. The shared-vocabulary assumption applies to the initial OPD setup, not to every OPRD-Bridge experiment.
2.2 The On-Policy Distillation Framework
At each training step, the student samples a response and the teacher is evaluated on the prefixes visited by that student. The canonical OPD objective is the expected sum of token-level reverse KL divergences between student and teacher next-token distributions.
- Student-generated prefixes reduce the training–inference mismatch associated with fixed teacher trajectories.
- Exact reverse KL requires a vocabulary-wide sum at every response position, motivating approximations for large vocabularies.
2.3 Three Output-Space Variants
Sampled-token OPD estimates reverse KL using the generated token’s student–teacher log-probability ratio; full-vocabulary OPD computes the exact vocabulary-wide divergence; top-k OPD restricts supervision to the student’s highest-probability tokens. These choices trade token-sampling variance, vocabulary-sized computation, and support truncation.
- Eq. (5) defines top-k KL using distributions renormalized on the selected support.
- The experimental setup instead specifies a probability-weighted sum of log-ratios over the top-16 tokens without that renormalization, leaving a discrepancy between the formal and evaluated objectives.
- Under the renormalized Eq. (5), singleton support has zero KL and therefore does not recover the sampled-token log-ratio objective, despite the stated analogy.
2.4 A Shared Structural Limitation
All three output-space variants compare next-token distributions without a separate target for intermediate hidden states. The paper argues that softmax invariance and weak sensitivity along some LM-head directions can make output agreement insufficient to ensure representation agreement.
- The clearest structural distinction is direct supervision at multiple intermediate layers rather than only at the final output.
- A tall LM-head matrix can have full column rank, so d being much smaller than the vocabulary size does not alone imply a nontrivial hidden-state null space.
3 On-Policy Representation Distillation
OPRD compares selected student and teacher hidden states before the LM head on the same on-policy sequences. Figure 2 distinguishes this hidden-space loss path from output-space reverse KL.
- Takeaways: OPRD supplies layer-wise targets, adds no token-sampling randomness to its fixed-rollout MSE gradient, and avoids vocabulary-sized tensors in the student distillation loss path.
- OPRD-Vanilla uses direct hidden-state matching; OPRD-Bridge compares representations through a frozen low-rank interface.
- OPRD can operate alone or be added to an OPD objective using the same rollout and teacher forward pass.
Both networks process the same student-generated sequence. OPD compares LM-head outputs, whereas OPRD compares selected intermediate representations and backpropagates only through the student.
3.1 The OPRD Objective
The OPRD objective averages squared student–teacher hidden-state differences over selected layers and valid supervised positions, with a 1/d normalization and stop-gradient on the teacher. Positions beyond a response’s valid length are masked out.
- The main OPRD-Vanilla configuration supervises all transformer layers and the last k = 2000 response tokens.
- The optional composite objective is L(θ) = LOPD(θ) + µLOPRD(θ), µ ≥ 0, as specified in Eq. (7).
- OPRD-only training is a separate configuration: setting µ = 0 in Eq. (7) gives OPD-only, although §4.1.1 labels that setting OPRD-only.
3.2 Why OPRD Works
Theorem 1 motivates OPRD through fixed-prefix gradient determinism, while Theorem 2 discusses softmax-invariant and weakly transmitted representation directions. These arguments require narrower interpretation than the paper’s claims of eliminating late-stage stochasticity and necessarily exposing nonzero null directions.
- For a fixed prefix, hidden-state MSE, full-vocabulary KL, and top-k KL on a fixed support are deterministic. Once the entire sampled rollout is fixed, the reused sampled-token estimator is also deterministic; its token-sampling variance is instead defined before conditioning on the current sampled token.
- The effective null space is NW = {∆h ∈ R^d : Whead∆h ∈ span{1}}. It describes perturbations that become constant logit shifts, but can contain only the zero vector.
- The accuracy plateaus do not directly measure gradient variance or establish that an LM-head null direction caused the residual accuracy gap.
3.3 OPRD-Bridge: Cross-Architecture Extension
3.3.1 The Shared Low-Rank Structure motivates the bridge using Qwen3-4B and Qwen3-1.7B representations collected on student responses to 2K DAPO-Math-17K prompts. After learning a student-side linear map into teacher PCA coordinates, rank 8 yields 95.0% average projected cosine similarity, while rank 2048 yields 77.0%.
- 3.3.2 Stage 1: Bridge Construction pairs layers by proportional spacing: ϕ(lS) = round(((lS − 1)/(LS − 1)) · (LT − 1)) + 1.
- The teacher projector is obtained from centered hidden-state covariance using at most 16,384 sampled token rows per layer. Both backbones remain frozen while student projectors are trained for 20 epochs.
- 3.3.3 Stage 2: Bridge-Based Distillation freezes both projector sets and updates only the student backbone. Projected vectors are L2-normalized before MSE, with supervision on the last 2000 response tokens; Algorithm 1 summarizes both stages.
Stage 1 trains student alignment maps against frozen teacher PCA targets while both backbones remain fixed. In Stage 2, the maps remain fixed and gradients pass through them into the student backbone.
3.3.4 Design Philosophy treats the projector pair as a fixed interface analogous to the shared vocabulary in logit distillation. The low-rank bridge is intended to retain well-aligned variation while excluding poorly matched architecture-specific directions.
- Figure 3 shows the change in trainable components between bridge construction and backbone distillation; Table 1 contrasts the representation interface with vocabulary-space alignment.
- Theorem 3 combines PCA reconstruction optimality with a modeled trade-off between discarded predictable signal and included prediction residuals. High aggregate cosine similarity does not establish that all task-relevant information is retained.
- The paper states that jointly training the student projector performs worse, but provides no quantitative frozen-versus-joint-training comparison.
4 Experiments
The experiments evaluate OPRD-Vanilla with an aligned same-architecture model pair and OPRD-Bridge with heterogeneous pairs. Figure 4 shows late-training separation between OPRD and the output-space baselines across AIME24, AIME25, and AIMO.
- The curves show raw Avg@16 measurements and a 5-step centered rolling mean.
- The visible separation is greatest on AIME24 and AIMO and smaller on AIME25; the curves include fluctuations rather than strictly monotonic improvement.
- The same-architecture experiments quantitatively compare sampled-token and top-16 OPD, not full-vocabulary OPD.
Late-training separation is clearest on AIME24 and AIMO, with a smaller difference on AIME25. The curves fluctuate and describe optimization outcomes; they do not directly measure gradient signal-to-noise ratios.
4.1 OPRD-Vanilla: Same-Architecture Distillation
4.1.1 Experimental Setup uses JustRL-Deepseek-1.5B as the frozen teacher and DeepSeek-R1-Distill-Qwen-1.5B as the student. They share the Qwen2.5-1.5B backbone with 28 transformer layers, hidden dimension 1536, approximately 151K vocabulary entries, and an aligned training lineage.
- Training uses DAPO-Math-17K, 8 prompts per global step, 2 student responses per prompt, sampling temperature 1.0, and maximum generation length 16,384 tokens.
- All methods receive 500 AdamW optimizer steps, peak learning rate 1×10−5, 3% linear warm-up, cosine decay, bf16, and FSDP on 8×A100 (80G). The reported micro-batch is B = 8.
- Evaluation uses temperature 0.7 and 16 responses per problem on AIME24 and AIME25, each with 30 problems, and AI-MO/aimo-validation-amc with 83 problems. Boxed answers are graded by exact match.
4.1.2 Main Results reports OPRD-Vanilla Avg@16 of 49.8/34.6/79.1 on AIME24/AIME25/AIMO, compared with teacher scores of 50.8/35.6/79.5. The stronger OPD baseline on each benchmark reaches 47.1/34.0/77.0, giving OPRD gains of 2.7/0.6/2.1 percentage points.
- The unmodified student scores 32.9/21.9/62.2; OPRD improves these by 16.9/12.7/16.9 points.
- OPD top-1 scores 42.3/33.5/77.0 and OPD top-16 scores 47.1/34.0/76.5, so increasing output support does not improve every benchmark.
- The authors describe the remaining teacher gaps as within evaluation noise, but Table 2 provides no confidence intervals or repeated-training-seed estimates.
OPRD exceeds both evaluated OPD variants on all three same-architecture benchmarks. Its teacher gaps are 1.0, 1.0, and 0.4 percentage points, but statistical equivalence cannot be assessed without uncertainty estimates.
4.1.3 Training Dynamics compares accuracy, rollout length, and representation cosine similarity. Figures 5 and 6 show OPRD responses settling around 5,700 tokens rather than approximately 7,000 for OPD, while supervised representation similarity approaches 1.0.
- The response-length diagnostic is response_length/mean, smoothed with window 15.
- The representation diagnostic is rep/cosine_similarity, smoothed with window 5; the plot includes an early dip before sustained improvement.
- These measurements show shorter generation and closer target alignment, but their co-evolution with accuracy does not establish a causal mechanism.
After early fluctuations, OPRD produces substantially shorter rollouts than either OPD run. Together with Table 2, this shows improved reported accuracy alongside fewer generated tokens in the aligned-model setting.
Supervised representation cosine similarity approaches 1.0 after an early dip. The curve demonstrates closer target alignment on measured positions, not a strictly monotonic trajectory or an isolated causal explanation of accuracy gains.
The composition experiment adds OPRD to sampled-token OPD with µ ∈ {0, 1, 10}. Figure 7 reports AIME24 Avg@16 of 42.3, 47.7, and 50.2, respectively.
- µ = 1 exceeds the standalone OPD top-16 result of 47.1; µ = 10 leaves a 0.6-point gap to the teacher’s 50.8.
- The sweep demonstrates complementarity for these three weights in this setup, not robustness across arbitrary weights or tasks.
AIME24 accuracy rises from 42.3 to 47.7 and 50.2 across tested OPRD weights of 0, 1, and 10. These three points demonstrate complementarity within the sweep, not a general monotonic relationship or broad tuning robustness.
4.1.4 Efficiency measures actor-update peak memory above its starting baseline, not total GPU memory. Table 3 reports 30.2 GB for OPD top-1, 45.0 GB for OPD top-16, and 20.5 GB for OPRD per GPU.
- The update_policy segment is instrumented with torch.cuda.reset_peak_memory_stats() and torch.cuda.max_memory_allocated(); the reported value is the maximum across ranks.
- The 500-step wall-clock times, excluding evaluation, are 813, 812, and 563 minutes. OPRD saves approximately 31% wall-clock time, equivalent to a 1.44× speed-up.
- Transient-memory reductions are 32% relative to top-1 and 54% relative to top-16. The teacher implementation still computes and discards full logits, so the savings do not mean all vocabulary-sized computation has been removed.
The savings concern actor-update transient memory above the starting baseline and 500-step training time excluding evaluation. OPRD uses 20.5 GB and 563 minutes, compared with 45.0 GB and 812 minutes for top-16 OPD.
4.1.5 Mechanistic Analysis motivates suffix supervision by comparing initial student–teacher last-layer cosine similarity at the beginning and end of responses. Figure 8 reports 91.65% for the last 50 tokens, 97.69% for the first 1600 tokens, and 95.42% over the full response.
- The prefix is initially more aligned than the suffix, providing an empirical rationale for the last-k heuristic.
- This diagnostic is not a downstream accuracy ablation comparing first-k, last-k, and all-position training.
Figure 9 tracks actor/pg_loss for the composite runs. Adding OPRD shifts a loss spike earlier, while all runs eventually approach approximately zero policy-gradient loss despite different final accuracies.
- The authors interpret the spike as a possible policy phase transition and leave its mechanism under investigation.
- Near-zero policy-gradient loss does not establish identical student and teacher distributions or prove that the remaining gap lies in an LM-head null space.
Figures 10 and 11 examine output-level changes under hidden-state supervision. OPD top-16 plus OPRD eventually achieves higher val-topk/overlap_ratio than OPD top-16 alone, while entropy changes occur earlier in the OPRD-augmented sampled-token runs.
- Top-16 overlap is the intersection size of student and teacher top-16 token sets divided by 16. The composite curve dips early before recovering above the baseline.
- actor/entropy and teacher/entropy are measured on the same changing rollout positions; teacher entropy can therefore drift despite frozen teacher parameters.
- The diagnostics describe temporally associated changes in policy behavior, but do not resolve the proposed reorganization mechanism.
The composite objective causes an early overlap dip before recovering above OPD alone. Representation supervision therefore changes output agreement, but the trajectory is not uniformly improving.
Entropy changes begin earlier with the larger tested OPRD contributions, and student and teacher traces later track closely. Teacher entropy is measured on changing student rollouts, so it is a moving distributional reference rather than a fixed target.
4.2 OPRD-Bridge: Cross-Architecture and Cross-Tokenizer Distillation
4.2.1 Experimental Setup pairs Qwen3-4B, with 36 layers and hidden dimension 2560, with Qwen3-1.7B-Base, with 28 layers and hidden dimension 2048. They share an approximately 151K vocabulary, allowing comparison against ordinary token-level OPD.
- Stage 1 uses 2000 DAPO-Math-17K prompts and 20 projector-training epochs; the default bridge rank is r = 8.
- Stage 2 supervises all 28 student layers and the last 2000 response tokens using L2-normalized projected MSE. It uses 2 responses per prompt, temperature 1.0, 8 prompts per step, maximum response length 16,384, AdamW with peak learning rate 1×10−5 and cosine decay, bf16, and FSDP on 8×A100 (80G).
- Evaluation reports Avg@16, Best@16, Dist-4g, and average response length. Results are selected from the checkpoint with the best average performance across the three evaluation benchmarks, rather than from a separately specified validation set.
4.2.2 Main Results shows an accuracy–generation trade-off. In Table 4, rank-8 OPRD-Bridge reaches average Avg@16 of 11.2%, below OPD top-1 at 15.3% but above OPD top-16 at 10.4%.
- OPRD-Bridge Avg@16 is 5.0/3.8/24.8 on AIME24/AIME25/AIMO, compared with OPD top-1 at 8.5/6.3/31.0.
- Its Best@16 is 20.0/13.3/59.0, matching the three reported OPD top-1 values and their average of 30.8%.
- Best@16 measures whether at least one of 16 attempts succeeds. It is not a general capability ceiling or a substitute for per-response reliability.
OPRD-Bridge matches sampled-token OPD’s reported Best@16 but has lower Avg@16. Subsequent OPD training raises average Best@16 to 33.9%, while average Avg@16 remains below standalone sampled-token OPD; the student AIMO Best@16 entry is also inconsistent with its Avg@16.
Table 5 shows substantial changes in generation behavior. OPRD-Bridge has average Dist-4g of 20.4 and average response length of 5333.1 tokens, compared with 9.4 and 16019.8 for OPD top-1, and 11.7 and 11410.2 for OPD top-16.
- The bridge’s benchmark-specific Dist-4g scores are 22.6/17.2/21.3, and its response lengths are 5909.4/5234.5/4855.3 tokens.
- Dist-4g is the unique-to-total 4-gram ratio. It measures lexical repetition, not reasoning correctness or diversity of solution strategies, and should be interpreted alongside the large length differences.
- Shorter responses and less repetition accompany lower Avg@16 than sampled-token OPD; the bridge does not dominate that baseline on average correctness.
The bridge has the highest unique-4-gram ratio and shortest responses among standalone distillation methods, alongside the accuracy trade-off in Table 4. Some OPD mean lengths exceed the stated training generation limit, but no separate evaluation limit is specified.
Continuing from an OPRD-Bridge checkpoint with OPD top-1 raises average Best@16 from 30.8% to 33.9% and Avg@16 from 11.2% to 13.6%. The paper proposes that output-level fine-tuning calibrates the mapping from the aligned backbone to token probabilities.
- The sequential pipeline’s average Dist-4g becomes 15.0 and average response length becomes 7168.5 tokens, between standalone bridge and standalone OPD behavior.
- Its average Avg@16 remains below standalone OPD top-1 at 15.3%.
- The additional training phase lacks a matched-total-budget control, and the LM-head calibration explanation is not isolated by a head-only ablation.
4.2.3 Ablation Studies varies bridge rank over r ∈ {1, 2, 4, 8, 16, 32, 64, 128, 256, 512, 1024, 2048}. Figure 12 shows projected cosine similarity peaking at 95.0% for r = 8, then declining to 94.1% at r = 32, 87.9% at r = 256, and 77.0% at r = 2048.
- Before projector training, similarity is at most 7.5%; after training, rank 1 reaches 72.3%.
- Ranks 4–16 form the reported high-alignment range of 93.7–95.0%.
- The experiment measures alignment rather than downstream accuracy by rank. It does not establish rank 8 as the optimal distillation rank or directly measure per-direction alignment decay.
Learned projected alignment peaks at rank 8 and declines as more teacher PCA directions are included. This supports a compact alignment interface, but does not determine the rank maximizing downstream accuracy or directly verify the assumed per-direction correlation threshold.
4.2.4 Cross-Tokenizer Distillation uses Phi-4-mini-reasoning as teacher and Qwen3-1.7B-Base as student, with bridge rank r = 4. The teacher has 32 layers, hidden dimension 3072, and an approximately 200K tiktoken-based vocabulary, while the student uses Qwen BPE and an approximately 151K vocabulary.
- Table 6 reports bridge Avg@16 of 4.0/1.0/10.0 and Best@16 of 13.3/13.3/37.3. The corresponding reported averages rise from 2.8% to 5.0% and from 4.5% to 21.3%.
- The OPD row says “collapses” without numeric results or a detailed cross-tokenizer output-alignment baseline.
- The method description does not specify how unequal teacher and student token boundaries are aligned for hidden-state matching, leaving an essential implementation detail unresolved.
The rank-4 bridge improves reported cross-tokenizer averages while remaining far below the teacher. OPD receives only a qualitative collapse label, and the student AIMO Best@16 value is below Avg@16, which is inconsistent under the stated definitions.
5 Discussion
The discussion treats representational alignment as the prerequisite for direct matching and the bridge as a way to construct a restricted shared interface. It proposes multi-model RL merging and on-policy self-distillation with privileged teacher information as applications.
- Neither multi-teacher merging nor privileged-information self-distillation is experimentally evaluated.
- Alignment-aware pre-training is proposed to reduce representation mismatch through shared initialization, representation regularization, or co-distillation; differing widths would still require a dimension adapter.
- These are proposed extensions, not established benefits of the reported experiments.
6 Related Work
The related-work section connects OPRD to output-space knowledge distillation, on-policy language-model distillation, and intermediate-feature matching. Its distinguishing combination is representation supervision on evolving student-generated autoregressive prefixes.
- MiniLLM [8] and GKD [1] motivate on-policy distribution matching; FitNets [28], TinyBERT [17], MobileBERT [33], and MiniLM/MiniLMv2 [34, 35] motivate feature-level transfer.
- The paper contrasts predictive-prefix targets with fixed-corpus feature matching and discusses auxiliary supervision, DINO [3], and representation engineering [49] as adjacent approaches.
- Its novelty claim concerns the combination of on-policy autoregressive sampling and teacher–student hidden-state distillation, not representation matching in general.
7 Conclusion and Future Work
The conclusion emphasizes near-teacher same-architecture mathematics accuracy with lower measured training cost and representation-space transfer across heterogeneous models. It leaves transfer beyond mathematical reasoning, cross-modal distillation, adaptive supervision, and the proposed phase-transition mechanism for future work.
- Other proposed directions include on-policy representation self-distillation, attention-map matching, alignment-aware pre-training, and representation-level diagnostics.
- The experiments use small model pairs and three long-chain-of-thought mathematics benchmarks; code generation, dialogue, agent interaction, larger-scale models, and multimodal settings are not tested.
- Explicit convergence-rate bounds and tighter spectral characterizations are identified as future theoretical work.
Appendix
- A Notation Summary provides two tables covering policies, architectural dimensions, response distributions, objectives, bridge projectors, and theoretical quantities. The notation distinguishes the mixing weight µ from layer-specific teacher centering means. Table 7 nevertheless describes µ as weighting OPD, whereas Eq. (7) weights OPRD, and lists a dimension adapter W attributed to §3.1 that is not defined there.
- B Formal Theoretical Guarantees, B.1 Setup and Assumptions, B.2 Variance of Sampled-Token OPD, B.3 OPRD Gradient Is Deterministic, and B.4 Signal-to-Noise Ratio Collapse of Sampled-Token OPD analyze fixed-prefix gradients under differentiability, finite-Fisher-information, bounded-log-ratio, and bounded-Jacobian assumptions. The fixed-prefix MSE gradient is deterministic, but Eq. (17) defines an unweighted score estimator while subsequent results use a log-ratio-weighted REINFORCE estimator; these are different estimators. The variance lower bound and asserted O(δ²) squared-gradient scaling are not established by the displayed arguments, and nonzero representation error does not by itself imply a nonzero parameter gradient.
- B.5 Formal Results for Theorem 2: LM-Head Information Bottleneck derives softmax invariance under constant logit shifts and discusses sensitivity along LM-head singular directions. The invariance identity holds for the stated effective null space, but that space need not contain nonzero vectors for a full-column-rank head. The quadratic KL argument does not establish the same bound for individual sampled log-ratios, which can be negative and vary to first order; the displayed spectral derivation yields a bound involving 1/σd² rather than the stated condition-number expression without additional normalization.
- B.6 Formal Results for Theorem 3: Optimality of the Low-Rank Bridge includes B.6.1 Setup, B.6.2 Rate–Distortion Optimality, and B.6.3 Bias–Variance Decomposition of Distillation Error. PCA minimizes teacher reconstruction MSE among rank-r orthogonal projections, while the modeled total error adds discarded predictable signal to included linear-prediction residuals under monotone per-direction alignment. Its marginal error is (\lambda_r(1-2\rho_r^2)), giving a threshold near (\rho_r^2=\frac{1}{2}), but this surrogate is not projected MSE alone, does not establish an optimal downstream rank, and is not a complete finite-rate rate–distortion construction.
Brief Thoughts
The strongest contribution is practical: an aligned teacher–student pair benefits from direct multi-layer supervision without a vocabulary-sized student distillation loss path. Tables 2 and 3 support that conclusion; the bridge results support a narrower claim of transfer with shorter, less repetitive outputs rather than superior average correctness.
The heterogeneous-model results need clearer reporting and stronger controls. Tables 4 and 6 report student AIMO Best@16 of 3.5% below Avg@16 of 7.5%, which is inconsistent under the stated metric definitions, while Table 5 reports some mean lengths above the 16,384-token training generation limit without specifying a different evaluation limit. Repeated training seeds, uncertainty intervals, independent checkpoint selection, matched-budget sequential baselines, and downstream rank ablations would make the comparisons more interpretable. Cross-tokenizer token-position alignment and quantitative frozen-versus-joint projector comparisons are needed to assess reproducibility and the bridge mechanism. The observed curves do not establish universal SNR collapse for output-space distillation or equate high projected cosine similarity with transferred reasoning capability.