Gorio Tech Blog search

Likelihood-Based Reward Designs for General LLM Reasoning | Summary

|

Contents

This article explains the key points of Likelihood-Based Reward Designs for General LLM Reasoning.

  • 2026-02-03 (arXiv)
  • Kwiatkowski, Ariel, Butt, Natasha, Labiad, Ismail, Kempe, Julia, Ollivier, Yann.
  • Meta FAIR, University of Amsterdam, New York University
  • Paper

Read this article in Korean


Summary

  • Likelihood-Based Reward Designs for General LLM Reasoning studies whether reference-answer likelihood can replace task-specific verifiers when training chain-of-thought (CoT) reasoning. It compares whole-answer probability rewards with log-probability rewards, including token-averaged and group-level variants.
  • Experiments use Llama-3.2-3B-Instruct and Qwen-2.5-3B-Instruct on MATH, DeepScaleR, Alpaca, and NuminaProof. Log-probability rewards achieve competitive greedy mathematical accuracy and substantially better reference-answer perplexity than Base RL or whole-answer probability rewards, but generally yield lower success under T = 1 answer sampling.
  • On long-form datasets, log-probability methods closely match supervised fine-tuning (SFT) in reference likelihood and perplexity, while whole-answer probability rewards do not produce comparable gains. This does not establish improved long-form reasoning: the learned CoTs become very short, making training effectively resemble SFT.
  • Minimum-length penalties, KL regularization, and warm-start training can preserve longer CoTs, but do not demonstrate an advantage over SFT within the reported compute budget. The evidence supports log-probability as a usable reference-based training signal across the tested domains, not as a demonstrated solution to learning useful explicit reasoning on non-verifiable tasks.

1 Introduction

Standard reasoning RL treats CoT tokens as actions and rewards the final answer using a correctness verifier. This applies naturally to short mathematical answers and executable code, but reliable task-specific checks are often unavailable for long proofs or open-ended responses.

  • Reference responses can remain available even when a reliable binary verifier is not.
  • The paper investigates probability and log-likelihood signals derived from those references, with particular emphasis on the connection between log-probability rewards and pretraining.

Our Approach and Contributions.

The contribution is a systematic comparison of likelihood-derived rewards across two model families and both short-answer and long-form settings. The study tracks answer success, reference-answer likelihood, perplexity, and CoT length to distinguish predictive performance from the development of explicit reasoning traces.

  • Log-probability variants improve reference likelihood across all four datasets; whole-answer probability rewards are ineffective on NuminaProof and unreliable on Alpaca.
  • Likelihood-based rewards do not require sampling the final answer for reward computation, although they still require generating a CoT and scoring the reference response.
  • The CoT-length analysis shows that successful likelihood optimization can coincide with largely eliminating explicit reasoning.

The paper distinguishes reference-grounded rewards from intrinsic rewards based on model confidence, entropy, or diversity, and from LLM-as-a-judge supervision. Its closest comparisons are VeriFree, JEPO, RLPR, and NOVER, which aggregate reference-answer probabilities in different ways.

  • VeriFree (Zhou et al., 2025) uses whole-answer probability. JEPO (Tang et al., 2025) takes the logarithm of a group-average answer probability rather than averaging individual CoT log-probabilities.
  • RLPR (Yu et al., 2025) motivates averaging individual token probabilities; NOVER (Liu et al., 2025) uses a geometric mean of per-token perplexities.
  • Reinforcement-pretraining (Dong et al., 2025) and LongForm (Gurung & Lapata, 2025) provide related continuation-based training approaches. Neither they nor NOVER are separate algorithms in the reported experimental comparison.

2 Method

The method separates the sampled reasoning trace from the reference answer used to evaluate that trace. The same reward construction can therefore score short exact answers and long reference continuations without requiring annotated CoTs.

Context: Chain-of-thought fine-tuning via Reinforcement Learning.

For prompt p, the policy πθ generates a CoT z and then an answer a. The conventional objective is Jθ = E[p ∼ D] E[z ∼ πθ(z|p), a ∼ πθ(a|p,z)] [R(z,a)], which can be optimized using Reinforce-style algorithms such as RLOO, GRPO, or PPO.

  • RLOO estimates an advantage by subtracting the mean reward of independently sampled traces for the same prompt, excluding the current trace.
  • The reward choice determines whether training emphasizes exact answer success or likelihood of the reference continuation.

RL fine-tuning with probability-based rewards.

With reference answer a⋆, the basic reward is R(z,a) = log πθ(a⋆|p,z). It can be computed by teacher-forcing the reference answer in one transformer pass and does not depend on a sampled final answer.

  • Average log-probability divides this reward by the reference-answer token count, changing the relative weighting of examples with different answer lengths.
  • VeriFree instead uses πθ(a⋆|p,z), the conditional expectation of an exact-match binary reward.
  • For long answers, the probability of matching the entire reference sequence can be extremely small. Taking the logarithm avoids compressing differences between these sequences into nearly indistinguishable zero rewards.

Because the reward depends on the current policy parameters, its gradient contains both a policy-gradient term for the CoT and a direct reference-answer likelihood gradient. Specifically, ∇Jθ = E[log πθ(a⋆|p,z) ∇log πθ(z|p) + ∇log πθ(a⋆|p,z)], so training is not merely RL against a fixed external scorer.

  • The first term reinforces traces that make the reference response more likely.
  • The second term is analogous to SFT conditioned on the sampled trace; when the trace largely disappears, training closely resembles direct answer SFT.

Algorithms and rewards tested.

The comparison includes SFT without a CoT, Base RL, Probability (VeriFree), Average prob (AvgProb), Log-prob, Average log-prob (AvgLogprob), and JEPO. Except for JEPO, RL advantages use a leave-one-out mean reward baseline for each prompt.

  • SFT predicts the reference answer directly from the prompt. Base RL generates both reasoning and an answer before applying the verifier.
  • AvgProb averages the conditional probabilities of reference tokens, whereas AvgLogprob averages their log-probabilities.
  • JEPO uses R(z1,…,zG) = log[(1/G) Σi πθ(a⋆|p,zi)] and subtracts a corresponding estimate excluding the current sample. Its experiments use G = 4 because larger groups are harder to implement efficiently.

Success metrics.

Evaluation distinguishes greedy answer success from T = 1 answer-sampling success and also measures reference-answer likelihood. The CoT-marginal answer probability is πCoT(a⋆|p) = Ez∼π(z|p)[π(a⋆|p,z)], so likelihood evaluation must account for the distribution of generated traces.

  • For the stochastic success-rate quantity, the method section reports estimating expected exact-match success by sampling a CoT and computing the conditional reference-answer probability. Greedy success instead evaluates deterministically decoded answers.
  • Per-token log-probability divides the summed answer log-probabilities by the total token count; per-answer log-probability normalizes each answer by its length before averaging. Perplexity is the exponential of minus per-answer log-probability.
  • Average CoT length includes formatting tokens, so a short nonzero trace need not contain substantive reasoning.

The paper estimates marginal log-likelihood with logprob-MC1 and, less frequently, logprob-MC32: it averages answer probabilities over N sampled CoTs before taking the logarithm. Concavity makes this estimator downward biased in expectation—not necessarily an underestimate for every realized sample—which matters when comparing CoT models with SFT.

  • Figures 18 and 19 show substantial MC1–MC32 differences for Base RL and whole-answer probability training on DeepScaleR.
  • Log-probability methods have smaller estimation gaps, and their likelihood advantage persists under MC32 in these comparisons.
  • Figure 18 labels its model as Llama 3.1 3B Instruct, whereas the experimental setup specifies Llama-3.2-3B-Instruct.

Averaging over 32 traces raises the marginal likelihood estimates especially for Base RL and Probability, but does not remove their disadvantage relative to log-probability methods. The caption's Llama 3.1 designation differs from the Llama-3.2 model specified in the experimental setup; that discrepancy cannot be resolved from the PDF.

Naive MC1 and marginal MC32 likelihood and perplexity curves on DeepScaleR, labeled Llama 3.1 3B Instruct in the source.
Naive MC1 and marginal MC32 likelihood and perplexity curves on DeepScaleR, labeled Llama 3.1 3B Instruct in the source.

3 Experimental Results

The experiments test whether verifier-free rewards preserve short-answer mathematical success and whether they improve prediction on long-form references. CoT-length measurements further assess whether those predictive gains accompany sustained explicit reasoning.

3.1 Setup: Datasets, Models, and Protocol

The two models are Llama-3.2-3B-Instruct and Qwen-2.5-3B-Instruct. MATH contains approximately 7,000 training problems after a random 10% validation holdout and is evaluated on its official test split; DeepScaleR contains approximately 39,000 training problems after a random 10% validation holdout and is evaluated on that held-out set.

  • Alpaca (cleaned) reserves 1,000 random examples for validation and leaves approximately 50,000 training samples.
  • NuminaProof filters theorem–proof items from NuminaMath-1.5, removes instances containing hyperlinks, sanitizes solutions, and reserves 1,000 validation examples, leaving approximately 50,000 training samples.
  • Mathematical training discards intermediate solutions and retains final answers, so dataset solutions do not supervise the generated CoT.

The setup compares SFT, Base RL, Probability, Log-prob, AvgLogprob, AvgProb, and JEPO. Verifiable experiments use G = 4 and G = 32, with JEPO restricted to G = 4, and a KL coefficient of 0.001; non-verifiable experiments use G = 4 without KL regularization in the main results. Verifiable result tables average over two seeds.

  • Training uses synchronous RLOO across 8 processes, AdamW with learning rate 10−5, a cosine schedule, 20 warm-up steps, and global gradient clipping at 1.0.
  • Each batch contains 8 questions with G CoTs per question. The default maximum completion length is T = 1024 tokens, and generation uses <think> and <answer> formatting.
  • The implemented Base RL reward is 100 for a correct parsed answer, 10 for an incorrect but correctly formatted answer, and 0 for an unparsable response—not the conceptual 0/1 reward. Mathematical training and evaluation use exact answer matching rather than semantic-equivalence verification.

3.2 Results on Verifiable Domains

At G = 32, likelihood-derived rewards achieve broadly comparable greedy mathematical success to Base RL. For Llama on MATH, Log-prob obtains 43.30 ± 0.10 versus 42.74 ± 0.66 for Base RL; on DeepScaleR, it obtains 32.40 ± 1.02 versus 28.55 ± 0.14.

  • For Qwen on MATH, Log-prob reaches 56.84 ± 0.21 versus 55.85 ± 0.46 for Base RL.
  • For Qwen on DeepScaleR, Log-prob is slightly below Base RL: 37.92 ± 0.11 versus 38.30 ± 2.40. The table therefore does not support a literal claim that every likelihood-derived variant improves every model–dataset combination.
  • No-CoT SFT has much lower greedy mathematical success, ranging from 9.88 to 18.32 across these four settings.

The four model–dataset blocks separate greedy success from stochastic success and reference likelihood. For Qwen on MATH, Log-prob reaches 56.84 ± 0.21 greedy success and perplexity 1.55 ± 0.03, while Base RL reaches 55.85 ± 0.46 and 8.25 ± 0.80; under T = 1 sampling, Base RL instead leads 55.36 ± 0.17 to 44.18 ± 0.42. Qwen DeepScaleR also qualifies the broad superiority claim, because Log-prob greedy success is slightly below Base RL.

Final verifiable-domain metrics for Llama and Qwen on MATH and DeepScaleR with G = 32, averaged over two seeds.
Final verifiable-domain metrics for Llama and Qwen on MATH and DeepScaleR with G = 32, averaged over two seeds.

The ranking changes under T = 1 answer sampling. For Qwen on MATH, Log-prob obtains 44.18 ± 0.42 sampled success versus 55.36 ± 0.17 for Base RL, despite slightly higher greedy success; its perplexity is 1.55 ± 0.03 versus 8.25 ± 0.80 for Base RL and 1.99 for SFT.

  • For Llama on MATH, Log-prob perplexity is 2.21 ± 0.06, compared with 13.87 ± 0.34 for Base RL, 7.14 ± 0.26 for Probability, and 2.63 for SFT.
  • Figure 1 shows similar final greedy success across RL variants, better likelihood under log-probability rewards, and an early CoT-length dip followed by recovery.
  • At G = 4, Table 3 shows no consistent greedy-success advantage for JEPO over simple Log-prob. Its likelihood ranking also depends on whether MC1 or MC32 is used.
  • The authors interpret the likelihood results as evidence of smoother answer distributions under log-probability training, but do not report a dedicated calibration metric.

The curves show similar final greedy success across RL variants but better reference likelihood under log-probability training. Its CoT length drops sharply early in training and then recovers, unlike the persistent collapse observed in the long-form experiments.

Training curves for Llama 3.2 3B Instruct on MATH with G = 32, showing success, likelihood, perplexity, and CoT length.
Training curves for Llama 3.2 3B Instruct on MATH with G = 32, showing success, likelihood, perplexity, and CoT length.

The smaller-group comparison adds JEPO and shows no consistent greedy-success advantage over simple Log-prob. For Llama on MATH, JEPO reaches 43.19 ± 0.10 versus 40.58 ± 2.63 for Log-prob, but on DeepScaleR it reaches 28.23 ± 0.11 versus 28.93 ± 1.55. Likelihood comparisons also require distinguishing MC1 from MC32, especially for JEPO.

Final verifiable-domain metrics with G = 4, including JEPO alongside the other reward variants and SFT.
Final verifiable-domain metrics with G = 4, including JEPO alongside the other reward variants and SFT.

3.3 Results on Non-verifiable Domains

On NuminaProof and Alpaca, Log-prob, AvgLogprob, and JEPO closely match SFT in reference-answer likelihood and perplexity. For Qwen on NuminaProof, Log-prob has per-answer average log-probability −1.0174 and perplexity 2.77, compared with −1.0172 and 2.77 for SFT; Probability remains at −1.1862 and 3.27.

  • For Llama on Alpaca, Log-prob and SFT both have perplexity 2.56, while Probability has 4.84, worse than the base model’s 3.85. Qwen Alpaca is a qualification: Probability improves from 4.04 to 3.66, but remains well above SFT at 2.44.
  • AvgProb generally approaches the log-probability family, although the paper describes its training as noisier. Token-average and whole-answer probability rewards therefore should not be assigned the same failure mode.
  • These results measure prediction of supplied references, not independently verified proof correctness or human-rated response quality.

Log-probability methods attain nearly the same long-form likelihood as SFT while producing very short traces. Qwen NuminaProof makes the distinction clear: Log-prob and SFT both have perplexity 2.77, with CoT lengths 5.3 and 5.0, while Probability retains 225.2 tokens and perplexity 3.27. Qwen Alpaca shows a modest Probability improvement over the base model, but not performance comparable to SFT.

Final non-verifiable-domain likelihood, perplexity, and CoT-length metrics on NuminaProof and Alpaca.
Final non-verifiable-domain likelihood, perplexity, and CoT-length metrics on NuminaProof and Alpaca.

3.4 Length of the Chain-of-Though During Training

Log-probability training initially shortens CoTs in both domain types, but the traces recover only in the verifiable experiments. On Qwen NuminaProof, Log-prob reduces average CoT length from 225.6 to 5.3 tokens, compared with 5.0 formatting tokens for no-CoT SFT; Probability retains 225.2 tokens without improving perplexity.

  • Figure 2 shows reference-likelihood improvement occurring alongside the rapid disappearance of explicit reasoning on NuminaProof.
  • Figures 20–22 compare pooled and per-question correlations using 100 random questions and 1000 CoTs per question at T = 1; only 20 questions are displayed.
  • On Numina, log-probability has global correlation 0.0486 but average local correlation −0.4662 with CoT length. On MATH, the corresponding values are −0.4511 and −0.3006.
  • MATH whole-answer probability instead has global correlation −0.0806 and average local correlation 0.0680. Local correlations are relevant to RLOO because its advantages compare traces for the same question.

On Qwen NuminaProof, improvements in reference likelihood accompany rapid CoT shortening and convergence near SFT. Whole-answer Probability remains near its initial likelihood while retaining long traces, showing that trace length alone does not establish a predictive benefit.

Training curves for Qwen 2.5 3B Instruct on NuminaProof, comparing per-answer log-probability, perplexity, and CoT length.
Training curves for Qwen 2.5 3B Instruct on NuminaProof, comparing per-answer log-probability, perplexity, and CoT length.

Numina has a slightly positive pooled correlation of 0.0486 but a negative average within-question correlation of −0.4662. Because RLOO compares traces within a prompt, the local relationship is consistent with pressure toward shorter traces even though the pooled statistic suggests almost no relationship.

Global and per-question correlations between reference-answer log-probability and CoT length for Llama 3B Instruct on Numina.
Global and per-question correlations between reference-answer log-probability and CoT length for Llama 3B Instruct on Numina.

On MATH, both the pooled correlation, −0.4511, and average within-question correlation, −0.3006, are negative. This is consistent with early trace shortening under log-probability rewards, but does not explain the subsequent recovery during mathematical training.

Global and per-question correlations between reference-answer log-probability and CoT length for Llama 3B Instruct on MATH.
Global and per-question correlations between reference-answer log-probability and CoT length for Llama 3B Instruct on MATH.

Whole-answer probability has average local correlation 0.0680 and pooled correlation −0.0806 with CoT length. Many traces receive near-zero probabilities, compressing reward differences that remain visible on the log-probability scale. The weakly positive local correlation is consistent with the absence of the same early shortening pattern under Probability training.

Global and per-question correlations between reference-answer probability and CoT length for Llama 3B Instruct on MATH.
Global and per-question correlations between reference-answer probability and CoT length for Llama 3B Instruct on MATH.

The authors test stronger KL regularization, penalties for traces below a length threshold, and an SFT warm start that trains reference answers after fixed generated CoTs while masking CoT gradients. These interventions preserve longer traces but do not establish better final reference-answer performance than SFT.

  • Figure 14 illustrates the length–performance trade-off on Llama NuminaProof: interventions preserving nontrivial CoTs generally slow or worsen likelihood improvement.
  • Figure 13 shows that warm-started CoTs stabilize around 100–200 tokens rather than approximately 5 tokens, but perplexity does not beat the compute-matched SFT baselines.
  • Tang et al. (2025) train JEPO approximately an order of magnitude longer, with a larger batch and lower learning rate. The present long-form result is therefore limited to the explored training budget.

Warm-start training partly prevents trace collapse, with lengths stabilizing around 100–200 tokens. The likelihood and perplexity curves nevertheless fail to surpass the compute-matched SFT reference levels, so preserving explicit traces does not by itself provide a long-form performance advantage.

Warm-started Llama 3.2 3B Instruct training on Numina, with log-probability variants and standard and warm-start SFT comparisons.
Warm-started Llama 3.2 3B Instruct training on Numina, with log-probability variants and standard and warm-start SFT comparisons.

The rows compare reward families while the columns track prediction quality and trace length. Length penalties maintain substantial CoTs, but generally slow likelihood improvement; stronger KL also worsens prediction quality in several reward families. The plot distinguishes mechanically retaining a trace from making that trace useful.

KL and minimum-length reward ablations for Llama 3.2 3B Instruct on NuminaProof.
KL and minimum-length reward ablations for Llama 3.2 3B Instruct on NuminaProof.

Discussion: Why don’t CoTs improve performance for non-verifiable domains? The paper proposes two explanations without resolving them: difficult credit assignment over long action sequences, and the possibility that long answers permit internal reasoning without a separate explicit CoT.

  • Credit-assignment difficulty could explain delayed recovery, but does not explain why verifiable Probability training lacks the same early dip or why length penalties fail to help.
  • Internal reasoning is a hypothesis drawn from related literature, not a mechanism measured in these experiments.
  • Negative local correlations support an initial pressure toward shorter traces, but do not identify why useful long-form CoTs fail to emerge later.

4 Conclusion

The conclusion presents log-probability rewards as a common reference-based objective for short verifiable answers and long non-verifiable continuations. In the tested settings, they combine competitive greedy mathematical success with improved perplexity and retain SFT-level long-form likelihood. Learning beneficial explicit CoTs on long-form tasks remains unresolved.

Appendix

  • A Losses and Advantages for the Rewards Considered derives the gradient for parameter-dependent rewards. Lemma 1 separates the score-function term from the direct reward derivative and shows that subtracting a baseline independent of the sampled trace leaves the expected gradient unchanged; stop-gradient operators implement the corresponding sampled surrogate.
  • B Experimental details specifies the optimizer, parallel execution, batch construction, answer extraction, dataset filtering, and prompting. Full Details on the Datasets. documents validation splits and removal of mathematical solutions; Prompting and formatting. specifies DeepSeek-R1–style tags, the 1024-token default cap, truncation at <answer for likelihood scoring, and the implemented 100/10/0 Base RL reward.
  • C Additional Experimental Results contains C.1 Verifiable Domains and C.2 Non-verifiable Domains. Figures 3–9 and Table 3 extend the mathematical comparisons across models and group sizes, while Figures 10–12 extend the long-form comparisons. They show the decoding-dependent mathematical trade-off and long-form likelihood approaching SFT as CoTs shorten.
  • D Attempted regularization methods adds (R_l(z)=r\cdot\min{ z -l_0,0}), with thresholds 100, 150, 300, and 500 tokens. The coefficient (r=\frac{\Delta R}{\Delta L}) is calibrated from the reward increase and length decrease over the first 40 training steps; Figures 14–17 examine this penalty and KL coefficients 0.001 and 0.1. Warm start. describes answer-only SFT after fixed generated CoTs, which partially stabilizes length without beating compute-matched SFT.
  • E Impact of the Marginal Log-Probability Estimate examines differences between MC1 and MC32 likelihood estimates. Figures 18 and 19 show larger estimator gaps for Base RL and Probability than for log-probability methods on DeepScaleR. Their likelihood performance remains worse after averaging over more traces.
  • F Correlation Analysis distinguishes pooled Pearson correlation from the average of within-question correlations. Figures 20–22 show that these statistics can differ in sign. The negative local relationship between log-probability and CoT length is consistent with the initial shortening pressure under group-relative RL.

Brief Thoughts

Table 1 shows why greedy accuracy alone is insufficient for comparing these objectives: similar mathematical accuracy can accompany substantially different reference likelihoods and stochastic success rates. The MC32 comparisons strengthen the likelihood finding in the settings where they are reported.

General applicability here means applicability of a reference-based objective, not demonstrated general reasoning improvement. The evidence comes from two 3B instruction-tuned models, four datasets, and the reported training budgets; long-form evaluation does not establish a benefit over SFT or independently assess proof correctness.

Base RL’s formatting reward and exact answer matching matter when interpreting the mathematical comparisons. The correlation study provides evidence consistent with early CoT shortening, but the cause of persistent long-form collapse remains unresolved.