Length Value Model: Scalable Value Pretraining for Token-Level Length Modeling | Summary
29 Apr 2026 | Paper Review Length-Controlled Generation Value Pretraining Output Length PredictionContents
- Summary
- 1. Introduction
- 2. Related Work
- 3. Length Value Model
- 4. Experiments: Validating LenVM as a Token-Level Length Signal
- 5. Qualitative Case Study: Length Tokens as Markers of Length Shifts
- 6. Ablations
- 7. Conclusion
- Appendix
- Brief Thoughts
This article explains the key points of Length Value Model: Scalable Value Pretraining for Token-Level Length Modeling.
- 2026-04-29 (arXiv)
- Zhang, Zhen, Yang, Changyi, Xia, Zijie, Yang, Zhen, Liu, Chengzhi, Weng, Zhaotiao, Liu, Yepeng, Chen, Haobo, Pan, Jin, Zhao, Chenyang, et al.
- University of California, Santa Barbara, Carnegie Mellon University, University of Wisconsin–Madison, LMSYS Org, Apple Inc.
- Paper
- Github
- Project Page
Summary
- Length Value Model: Scalable Value Pretraining for Token-Level Length Modeling introduces LenVM, a prefix-conditioned estimator of discounted remaining generation length. It trains a scalar value head with dense Monte Carlo targets constructed automatically from sampled completions, without additional human annotations or a reward model.
- LenVM improves Qwen2.5-7B-Instruct’s Equal-To Length Score from 30.9 to 64.8 on LIFEBench-token; combining it with eight-call LCG raises the score from 71.3 to 83.6 and reduces mean absolute relative deviation from 14.0% to 5.8%. Prediction accuracy and validation loss improve with model and data scale, while guided decoding achieves higher task accuracy than hard truncation at similar average output lengths.
- LenVM estimates a policy-conditioned discounted horizon, not the conditional mean of raw token count, and optimized guided decoding still incurs 1.90× measured latency. The stated exponential-tilting formula with β < 0 favors more negative values, which the paper’s return definition associates with longer—not shorter—continuations; the implemented sign convention therefore needs clarification.
1. Introduction
Generation length affects decoding compute, KV-cache growth, latency, and the amount of available reasoning. The paper motivates a prefix-dependent horizon estimator that updates after individual token choices, rather than relying solely on sequence-level penalties, prompt instructions, or pre-decode forecasts.
- LenVM evaluates distance to termination rather than task utility, mapping widely varying completion lengths into a bounded value space.
- Its supervision is annotation-free, dense, and unbiased for the conditional discounted-return target under a fixed rollout policy. This does not imply unbiased raw-length prediction or calibration under arbitrary policy changes.
- The experiments evaluate inference-time applications. PPO-style use is developed conceptually, with empirical RL fine-tuning left to future work.
2. Related Work
The related-work discussion covers length-constrained generation, static and progressive output-length prediction, and value-based formulations of generation cost. LenVM uses a separately trained token-level estimator with discounted-return targets derived from realized completion lengths.
2.1. Length-Controlled Generation
Prior length-control methods include countdown prompting, Plan-and-Write scaffolds, iterative Metropolis–Hastings sampling, length-difference positional encoding, and hidden length-tracking tokens. LenVM instead scores candidate next states during decoding, leaving the base generator’s parameters unchanged and allowing combination with an outer iterative controller such as LCG.
2.2. Output Length Prediction
The paper distinguishes prompt-boundary prediction from progressive remaining-length prediction. It compares against entropy-guided token pooling (EGTP) and progressive length prediction (PLP), while also discussing forecasts from frozen layerwise representations and horizon prediction for code infilling.
- Length prediction could inform scheduling, batching, and memory planning, but the paper evaluates prediction metrics rather than end-to-end scheduler improvements.
- Unlike predictors trained directly on raw token counts, LenVM learns a bounded discounted-horizon target.
2.3. RL Framing and Reward Shaping of Generation Length
A constant per-token cost makes remaining generation horizon a value-function quantity. The paper relates this formulation to PPO, GAE, value pretraining, and adaptive length penalties, while distinguishing direct quality–efficiency optimization from potential-based shaping that preserves the original objective under fixed-potential assumptions.
3. Length Value Model
LenVM treats autoregressive decoding as an episodic process and produces a scalar estimate from each prefix, starting at the last prompt token. Figure 1 connects sampled completions, automatically computed discounted returns, final-layer representations, and token-wise regression.
- Additional prompts and completions expand supervision without additional labeling.
- The project page is https://github.com/eric-ai-lab/Length-Value-Model.
Every sampled completion supplies a remaining-horizon target at each nonterminal prefix. The value head maps final-layer representations to bounded predictions, while model size, prompt count, and completions per prompt provide separate scaling axes.
3.1. Modeling Length as a Discounted Return
For a completion of length L and a state after t generated tokens, the nonterminal reward is r_t = −(1 − γ), with r_L = 0. The target is (G_t=-(1-\gamma^{L-t})), satisfying G_t = r_t + γG_(t+1): shorter remaining horizons approach 0, whereas longer horizons approach −1.
- For a fixed rollout policy π, the value is V^π(s_t) = E_π[G_t | s_t]. A realized return can be inverted exactly as l = log(1 + G_t) / log γ.
- Inverting a conditional mean return does not generally recover the conditional mean token count; Appendix G proves underestimation when remaining length is stochastic.
- Larger γ preserves resolution over longer horizons, while smaller γ concentrates resolution nearer termination.
3.2. LenVM Head
The LenVM head is a two-layer MLP with SiLU activation applied to the final-layer hidden state h_t. Its output is (V_\theta(s_t)=-\sigma!\left(W_2\operatorname{SiLU}(W_1h_t+b_1)+b_2\right)), constraining predictions to (−1, 0).
3.3. Training Objective
Training minimizes squared prediction error summed over all nonterminal states in the minibatch and divided by the total number of generated tokens. Completions come from a fixed generator checkpoint and decoding policy, and each target is computed exactly from the realized remaining length.
- Token-uniform averaging preserves the state-conditional return target. Averaging token losses within each sequence before averaging sequences introduces weights proportional to 1/L and generally changes that target.
- GAE provides an optional bootstrapped variant, but λ = 1 performs best in the reported experiments. Complete trajectories already supply dense full-return targets.
4. Experiments: Validating LenVM as a Token-Level Length Signal
The experiments evaluate explicit length control, continuous accuracy–length steering, prompt-boundary and progressive prediction, and scaling of the value-learning objective. Separate latency measurements quantify the added forward-pass cost rather than treating fewer output tokens as equivalent to faster inference.
4.1. Experimental Setup
General LenVMs use a mixture of OpenCodeReasoning-2 (Python; 1.42M), WildChat (529k), and DeepMath-103K (103k), listed in Table 1. Qwen2.5 experiments initialize LenVM from the Qwen2.5-Instruct family, including vision-language checkpoints; Qwen3 experiments use Qwen3-Base initialization.
- Appendix B specifies 2 epochs, learning rate 2 × 10^−5, batch size 1024, BF16, and up to 16 completions per prompt sampled with temperature 1.0 and top-p 1.0 unless otherwise stated.
- Reported discount factors are 0.997 for Qwen2.5-Instruct and Qwen2.5-VL-Instruct, and 0.9998 for Qwen3-Instruct. The latter naming differs from the main text’s Qwen3-Base initialization description.
- Ablation and scaling runs use 100k examples per dataset except math, which uses all 95k available examples; control and efficiency experiments use the full training sets. Training validation uses 8k randomly sampled examples.
The training mixture covers code, instruction following, and mathematics, with unequal listed corpus sizes. These source scales do not establish equal domain weighting or identical sample counts in every experiment.
4.2. Application 1: Length-Controlled Generation
LIFEBench-token evaluates Equal To, At Most, and At Least token-count constraints on 360 bilingual base instances; Section 4.2 reports targets from 32 to 1024 tokens. LenVM scores candidate next states and selects the closest target value for Equal To, the most negative value for At Least, or the value nearest 0 for At Most.
- With temperature 1.0, top-p 0.999, and top-k 15, Qwen2.5-7B-Instruct improves from 30.9 to 64.8 Equal-To LS and from 71% to 44% Equal-To LD. At-Most LS decreases from 98.5 to 96.1, so improvements are not uniform across constraints.
- Table 2 compares one-pass methods with multi-call controllers. One LenVM pass exceeds Prompt Best-of-8 in LS, but a generation call is not an equal-compute unit because LenVM adds candidate scoring.
- Table 6 shows complementary gains when inserting LenVM inside LCG proposals: at 2, 4, and 8 calls, LS rises from 60.6, 69.0, and 71.3 to 73.8, 80.5, and 83.6, respectively.
- Appendix A describes the original ten-target grid from 16 to 8192 and says it is preserved, leaving the evaluated target range unclear. Token-based scores are not directly comparable with the original word/character leaderboard.
Equal-To adherence improves substantially, including Qwen2.5-7B-Instruct’s LS increase from 30.9 to 64.8. Its At-Most score decreases, and differing call budgets prevent interpreting the table as either uniform improvement or an equal-compute ranking.
Adding LenVM improves LCG at every tested outer call budget, with the eight-call combination reaching 5.8% LD and 83.6 LS. Matching rounds controls the revision budget, not total inference compute, because LenVM adds candidate scoring within each call.
4.3. Application 2: Performance–Efficiency Trade-off
The paper defines value-guided exponential tilting as (p’(x)=\frac{p(x)\exp(\beta\hat v(x))}{\sum_{x’}p(x’)\exp(\beta\hat v(x’))}), with β < 0. Figure 2 reports better Pass@1–length frontiers than hard truncation on GSM8K, MATH500, and MathVista; near 200 tokens on GSM8K with Qwen2.5-3B-Instruct, reported Pass@1 is about 63% versus 6%.
- These experiments sample 64 completions per question with temperature 1.0 and top-p 1.0; guided decoding uses min-p 0.01 to limit candidate batches. Hard-budget outputs exceeding B are truncated and marked incorrect.
- Nearest-length comparisons in Table 5 use Qwen2.5-7B-Instruct. On MATH500, LenVM obtains 45.59 Pass@1 at 306 tokens versus EOS calibration’s 28.56 at 307 tokens, and 47.50 at 317 tokens versus budget prompting’s 39.66 at 333 tokens.
- Table 7 reports Base + LenVM at 54.17 Pass@1 and 5611 tokens on AIME2025, compared with ART at 51.25 and 5599 tokens. Steering ART further reduces length with an accuracy cost, demonstrating compatibility rather than uniform superiority over RLVR training.
- The displayed formula conflicts with the stated shortening interpretation: β < 0 increases relative probability for more negative v̂, although Section 3 associates these values with longer horizons. The reported curves do not establish which sign or value convention was implemented.
The reported guided-decoding curves retain higher accuracy than hard truncation at similar average lengths on all three benchmarks. The roughly 63% versus 6% GSM8K comparison is substantial, but truncation is a restrictive baseline and the displayed steering convention has an unresolved sign mismatch.
Nearest-length comparisons strengthen the evidence beyond hard truncation. On MATH500, LenVM’s 45.59 Pass@1 at 306 tokens exceeds EOS calibration’s 28.56 at 307 tokens; budget-prompt comparisons also favor LenVM, although paired average lengths are not always identical.
Base + LenVM achieves 54.17 Pass@1 at 5611 tokens, compared with ART’s 51.25 at 5599 tokens. Applying LenVM to ART further shortens responses while reducing accuracy, demonstrating compatible decoding-time control rather than dominance across training protocols or operating points.
4.4. Application 3: Generation Length Prediction
For prompt-boundary evaluation across scales and domains, each held-out prompt receives N = 64 completions. The reference horizon is obtained by averaging u(L) = 1 − γ^L and then inverting that average, so the reported MRE measures a return-consistent horizon rather than mean raw completion length.
- Table 3(a) reports MRE reductions from 17.0%, 29.0%, and 33.0% at 1.5B to 9.8%, 14.9%, and 17.1% at 32B for math, code, and instruction following.
- Table 3(b) compares predictors with a frozen Qwen2.5-7B-Instruct generator on 8192 held-out DeepMath-103K prompts and approximately 108,000 prefix states. LenVM reduces prompt MRE from 32.56 to 26.90 and progressive MRE from 26.71 to 15.10.
- Progressive Within-128 accuracy increases from PLP’s 30.74 to LenVM’s 60.70. Appendix B specifies realized completion lengths as the common targets for this specialized comparison, distinct from the 64-sample transformed reference in the scaling evaluation.
Prompt-boundary MRE decreases in all three domains as model size increases. The specialized-baseline panel separately shows improvements in error, rank correlation, and Within-128 accuracy, particularly for progressive prediction; its realized-length evaluation targets differ from the transformed references in the scaling panel.
4.5. Scalability of LenVM
Figure 3 varies model size from 0.5B to 32B, training-question count across 10k, 30k, and 100k, and samples per question across 1, 2, 4, and 16. Larger models, more prompts, and more sampled trajectories generally yield lower validation loss within these tested ranges.
- The curves show empirical scaling trends, not a fitted scaling law or a compute-optimal allocation between model size, prompt coverage, and repeated sampling.
Larger models, broader prompt coverage, and more completions per prompt generally reach lower validation losses. These comparisons do not quantify compute-optimal allocation or the cost-normalized benefit of each scaling direction.
4.6. Inference Cost
Candidate next states are evaluated in one packed SGLang call with shared-prefix attention and cached LenVM KV states. Skipping value scoring when generator entropy is below 0.5 reduces active steps to 19.5% and measured end-to-end latency from 4.55× to 1.90× vanilla decoding, as reported in Table 4.
- The latency setup uses Qwen2.5-7B-Instruct as generator and Qwen2.5-1.5B-Instruct initialization for LenVM. The remaining overhead depends on hardware and serving load.
- Table 9 reports packed forward times of 32.18 ms, 34.16 ms, and 33.49 ms for K = 2, 3, and 5. This supports packed scoring rather than K serial calls, but not runtime independence from K outside that narrow range.
- The paper does not establish wall-clock savings from shorter generated responses.
Entropy skipping reduces value-model activation from every decoding step to 19.5% of steps and lowers relative latency from 4.55× to 1.90×. The optimized configuration therefore still takes 1.90 times the vanilla decoding time in this setup.
Packed scoring takes between 32.18 and 34.16 ms across K = 2, 3, and 5, supporting shared-prefix batching rather than K independent calls. The narrow tested range does not establish constant cost for larger candidate sets.
5. Qualitative Case Study: Length Tokens as Markers of Length Shifts
The case study computes the residual s_t = r_(t−1) + γV_t − V_(t−1) on 8k DeepMath-103K examples and groups frequent tokens above 0.01 or below −0.01. Figure 4 contrasts reasoning-pivot tokens such as wait, think, and try with closure-associated tokens such as therefore, clearly, and perfect.
- These groups describe associations with changes in predicted horizon, not causal effects of individual tokens.
- The paper labels positive residuals as longer-horizon shifts and negative residuals as shorter-horizon shifts. Under its negative-cost convention, a positive residual instead indicates a next-state value above the Bellman-predicted value, corresponding to a shorter proxy horizon relative to that prediction.
The word clouds display frequent tokens selected by positive and negative residual thresholds on mathematical trajectories. Reasoning pivots and closure language appear in different groups, but these associations do not establish causation, and the horizon-direction labels conflict with the negative-cost residual convention.
6. Ablations
The ablations examine target representation, batch construction, discount factor, numerical precision, and candidate-set size. Unless otherwise specified, the model, optimizer, and evaluation protocol are held fixed.
6.1. Length-Space Representation
Figure 5(a) compares Length + Softplus, Normalized Length + Sigmoid, Log Length + Softplus, and Discount Return + Sigmoid. Discounted returns achieve the lowest plotted average absolute length error; log-length regression improves substantially over maximum-normalized length but does not match discounted returns.
- The paper attributes poor normalized-length performance to compression of common short lengths into a narrow region and notes that discounted targets satisfy a stepwise Bellman recursion.
- The comparison does not isolate Bellman alignment from numerical conditioning as the cause of the advantage. Log-length and raw-length errors are also close late in training, limiting a blanket ranking of those two alternatives.
6.2. Batch Construction Strategy
Figure 5(b) compares fully shuffled training with batches that group multiple completions from the same prompt, holding the amount of data per update fixed. Shuffling produces lower evaluation loss in the tested setting.
6.3. Discount Factor γ
Figure 6 compares γ = 0.99, 0.995, and 0.999 at 0%, 25%, 50%, and 75% of generation. Larger γ performs better early, while smaller γ performs better near termination; the proposed selection rule sets 1 − γ^L_0.99 = 0.99 using the 99th-percentile generation length.
- Section 6.3 says larger γ compresses long horizons more aggressively, whereas Section 3.1 says it preserves more long-horizon resolution. The position-dependent empirical results can be reported without adopting this inconsistent explanation.
The panels show a position-dependent trade-off: γ = 0.999 performs best near the prompt boundary, while γ = 0.99 performs better late in generation. Intermediate positions have smaller separations, supporting a qualified trend rather than a universal ranking.
6.4. Numerical Precision
Figure 5(c) shows similar final losses for fp16, bf16, and fp32, despite differences early in training. Appendix E and Figure 7 analyze a local logit-perturbation proxy that approaches 1/k at long horizons and has larger relative effects near one-token horizons.
- The analysis uses k = 128 for bfloat16, 1024 for fp16, and 8388608 for fp32. It is a first-order representational-resolution proxy, not a bound on all inference or training errors.
The target-representation panel favors discounted returns, the batching panel favors shuffling, and the precision panel converges to similar final losses despite early differences. The first panel measures absolute length error, whereas the other panels report test loss.
The two views translate a local logit-perturbation proxy into horizon-dependent relative resolution. For the illustrated bfloat16 proxy, long-horizon sensitivity approaches 1/k = 0.0078, while one-token horizons have larger relative sensitivity; the curves are analytical approximations, not measured prediction errors.
6.5. Candidate-Set Size
Table 8 shows that increasing K from 1 to 128 improves LIFEBench-token LS from 29.3 to 72.1 and reduces LD from 100.2% to 18.5%, although intermediate results are not strictly monotone. For MATH500 with β = −100, the largest length reduction occurs from K = 1 to K = 2; accuracy and average length change less across larger tested sets.
- Hard target matching gains additional alternatives from a larger candidate pool, whereas exponential tilting retains the generator’s probability weighting.
- These K-sweep measurements are a separate ablation and should not replace Table 2’s control results.
Hard length matching benefits overall from a larger candidate pool, although control worsens from K = 2 to K = 8. For MATH500 steering, the largest length change occurs from K = 1 to K = 2, with smaller changes across subsequent tested settings.
7. Conclusion
The control, prediction, and scaling experiments support discounted remaining horizon as a useful token-level regression target. The evidence concerns inference-time utility and annotation-free supervision; empirical RL gains, calibration under unrelated generators, and net serving-efficiency gains remain unestablished.
Appendix
- A. LIFEBench-token Evaluation Details describes QA, summarization, reasoning, and creative generation in English and Chinese, with 360 base instances and an original ten-target grid producing 10,800 constraint instances. LD is signed relative length error per example, aggregated as mean absolute LD for Equal To. LS lies in [0, 100], uses exponential penalties with k_1 = 5 for under-generation and k_2 = 2 for over-generation where a constraint is violated, and assigns 100 when it is satisfied; prompts combine the task instruction with a natural-language token requirement.
- B. Additional Experimental Details specifies training, sampling, tokenizers, and baseline adaptations, including CAPEL countdown-markup removal, segmented Dynamic Feedback, and LCG proposal acceptance. LCG + LenVM preserves the outer controller and applies Equal-To guidance within each proposal using K = 15. Prediction baselines share realized-length targets, while policy-dependence and latency protocols clarify that generator-anchored candidates limit—but do not eliminate—off-policy mismatch.
- C. Additional Experimental Results reports packed end-to-end decoding latency, nearest-length quality comparisons, and the full two-, four-, and eight-call control comparison. It also compares the base model and ART on AIME2025, demonstrates further shortening of ART with an accuracy cost, and provides candidate-size sensitivity and packed scoring times. Generation-call budgets, output token counts, and measured latency represent different costs and should not be treated as interchangeable.
- D. LenVM as a Length-Specific Value Function in RL defines a discounted length objective, separate task and length critics, and an additive combination of their GAE advantages. Adding a length advantage changes the optimized objective; using LenVM as a fixed potential with matched discounting and zero terminal potential instead yields a telescoping shaping term that preserves return ordering from the same initial state. The appendix argues for dense credit assignment, a matched baseline, and separation of quality from efficiency over terminal length penalties, but reports no empirical RL gains and leaves these uses to future work.
- E. Finite-Precision Analysis of Relative Length Resolution rewrites the inverse horizon as l̂ = ln(σ(−z)) / ln γ and models local logit perturbation as |δz| ≈ |z|/k. For the modeled perturbation, relative length sensitivity approaches 1/k at long horizons. This analysis addresses representational resolution for horizons from 1 to 32k and complements the empirical precision ablation rather than bounding total numerical error.
- F. Future-Dependent Weighting Changes the Regression Target derives the weighted optimum V*_w(s) = E[wG | s] / E[w | s], where the conditional weight expectation is positive. Constant or positive current-state-only weights preserve the per-state conditional mean, whereas future-dependent weights generally shift it. Trajectory averaging supplies the example w = 1/L, downweighting longer completions and motivating token-uniform regression.
- G. Why Inverting the Transformed Horizon Underestimates Expected Remaining Length uses the strict concavity of u(L) = 1 − γ^L and Jensen’s inequality to show u^−1(E[u(L)]) ≤ E[L], assuming the expectations exist. Equality requires deterministic remaining length. Thus, even exact conditional-mean prediction in value space does not produce unbiased mean-token-count prediction after inversion.
- H. Derivation of the Exponential Tilting Solution in the Performance–Efficiency Trade-off Experiment derives the displayed Gibbs distribution from the KL-regularized objective for β < 0. Its algebra matches that objective, but interpreting lower values as shorter horizons contradicts Section 3’s negative-cost convention. The additional claim that β > 0 makes the objective unbounded below is also incorrect on a finite candidate simplex with positive base probabilities: the objective remains bounded, although the convex-minimization argument no longer applies.
Brief Thoughts
The main contribution is the reusable supervision construction: completion lengths supply dense, policy-conditioned value targets without annotations, and one estimator supports several inference-time applications. The nearest-length MATH500 comparisons and prediction gains provide stronger evidence of utility than hard truncation alone. The steering-sign mismatch is the main reproducibility concern. The TD-residual labels and discount-factor explanation also require clarification before the equations can serve as an executable decoding specification. Length Score measures graded adherence, not exact-match success. One-pass LenVM still has 44% mean absolute relative deviation in the Qwen2.5-7B comparison. The 1.90× latency result rules out a blanket speed claim. Policy-shift calibration, an unambiguous benchmark target range, and empirical RL evaluation would clarify the method’s scope.