Reinforcement Learning with Metacognitive Feedback Elicits Faithful Uncertainty Expression in LLMs | Summary
30 Jun 2026 | Paper Review Faithful Calibration Reinforcement Learning Uncertainty Expression Metacognitive FeedbackContents
- Summary
- 1 Introduction
- 2 Related Work
- 3 Method
- 4 Experimental Setup
- 5 Results
- 6 Conclusion
- Appendix
- Brief Thoughts
This article explains the key points of Reinforcement Learning with Metacognitive Feedback Elicits Faithful Uncertainty Expression in LLMs.
- 2026-06-30 (arXiv), Preprint
- Liu, Gabrielle Kaili-May, Caciularu, Avi, Yona, Gal, Szpektor, Idan, Cohan, Arman.
- Yale University, Google Research
- Paper
- Github
Summary
- The paper studies faithful calibration (FC): agreement between an LLM’s expressed uncertainty and intrinsic confidence estimated from sampled-response consistency, rather than agreement between confidence and answer correctness. It trains sentence-level numerical confidence using reinforcement learning with metacognitive feedback (RLMF) and metacognitive data selection, then edits the outputs to express uncertainty linguistically.
- On 10 English evaluation datasets, RLMF improves average FC over the reported baselines, including an improvement of up to 63% over standard RL in the numerical setting. A separate human comparison favors its rewritten responses over FUT on diversity, naturalness, helpfulness, and contextual suitability.
1 Introduction
Factual calibration can conceal a mismatch between what a model communicates and what its sampled responses suggest it would say consistently. The authors address this mismatch by training numerical confidence scores before mapping them to context-sensitive hedges, so the rewriting stage can change presentation without repeating RL training.
The workflow separates sampled completions, rewards, self-judgment, and advantage scaling from the subsequent conversion of confidence scores into hedged prose. Editing follows numerical FC training rather than supplying its training reward.
2 Related Work
Prior internal-feedback RL has used signals including confidence and entropy; RLMF instead uses the accuracy of a model’s judgment of its own FC performance to adjust its learning signal. Earlier linguistic FC work uses prompting, steering, or supervised fine-tuning, but the paper argues that these approaches leave numerical FC and natural, varied long-form hedging insufficiently addressed.
3 Method
The framework separates confidence training from linguistic editing. Stage 1 selects training examples and applies RLMF to sentence–confidence outputs; Stage 2 maps those scores to hedges and instructs an editor to preserve the answer’s factual claims.
3.1 Reinforcement Learning Setup
For each query, GRPO samples G completions with sentence-level confidence scores. Its weighted reward covers confidence faithfulness, factual calibration, correctness, and strict and soft format adherence; the reported implementation computes advantages by subtracting the group mean reward without dividing by its standard deviation.
- The main reward weights are 12 for faithfulness, 1 each for factual calibration and correctness, and 3 each for the two format rewards.
The diagram places the comparison of predicted and gold FC performance between reward computation and the GRPO policy update. It shows metacognitive quality modifying the faithfulness component of an advantage for above-average-faithfulness completions, rather than replacing the other rewards.
3.2 Reinforcement Learning with Metacognitive Feedback (RLMF)
For each completion, the current policy predicts an FC score; the gold FC level is the fraction of sentences whose expressed and sampling-estimated intrinsic confidences differ by less than τ = 0.10. RLMF sets Zg = 1 − (Fpred − Fgold)² and, only when a completion’s weighted faithfulness exceeds the group mean, multiplies that component of its advantage by k + Zg with k = 1; the other reward components remain unscaled.
- The condition keeps self-assessment accuracy subordinate to above-average performance on the primary faithfulness objective.
3.3 Metacognitive Data Selection
For each candidate training example, the model produces a linguistically hedged response and rates from 0–100 how closely its linguistic decisiveness matches its internal confidence. The method takes half of the target training examples from each end of this self-rating distribution, rather than selecting only low-rated examples.
3.4 Rewriting Protocol
The editor receives the original answer, candidate hedges associated with each sentence’s numerical confidence, and a target style or user context. The score–hedge map groups expressions by mean human-rated confidence into 0.05-wide bins; the main single-pass procedure samples 20 hedges from the relevant bin and requests fluent editing without changing factual claims.
4 Experimental Setup
Experiments use Llama3.1-Instruct (8B) and Qwen3 (1.7B, 4B, 8B), evaluating 1,000 test examples from each of 10 English datasets. After pre-SFT for the output format, the main RLMF setup trains on 2,000 metacognitively selected PopQA examples; alternative training datasets are SelfAware, HaluEval, and UMWP. Comparisons include MetaFaith, FUT, standard RL, and prompted proprietary models.
FC is measured with cMFG*, which uses equal-mass intrinsic-confidence bins weighted by their widths over the model’s observed confidence range. Sentence-level intrinsic confidence is estimated through NLI consistency judgments against 20 additional sampled responses; accuracy is judged by an LLM, while Brier Score (BS) measures alignment between intrinsic confidence and correctness.
For the two 8B open models, the RLMF rows have higher average cMFG* than their respective standard-RL rows. Separate numerical and rewritten linguistic rows allow the effect of editing to be inspected, while accuracy and BS expose outcomes beyond FC; the prior-method comparisons concern linguistic FC.
5 Results
In Table 1, Llama3.1-8B-Instruct obtains average numerical cMFG* of 0.84 with RLMF, versus 0.77 with standard RL; rewriting yields linguistic cMFG* of 0.82, versus 0.67 for MetaFaith and 0.66 for FUT. Qwen3-8B obtains 0.83 with RLMF and 0.83 after rewriting, versus 0.51 with standard RL; its accuracy is 0.57 with RLMF versus 0.55 for the original model.
On PopQA, the original Llama3.1-8B-Instruct model and FUT express relatively high confidence in low-intrinsic-confidence bins. The RLMF-trained model’s expressed-confidence curve is closer to intrinsic confidence across the displayed bins, with higher faithfulness scores.
Each listed RLMF training dataset improves average cross-task numerical cMFG* over the original models. PopQA gives the highest listed scores, while SelfAware, HaluEval, and UMWP test sensitivity to the training source rather than only performance on that source.
Metacognitive selection has the highest cMFG* for both 8B models among the three tested selection strategies. Accuracy and BS do not always follow the same ordering: for Llama3.1-8B-Instruct, random selection has a lower BS than metacognitive selection.
Training on SelfAware, HaluEval, or UMWP instead of PopQA gives cross-task numerical cMFG* of 0.80–0.81 for Llama3.1-8B-Instruct and 0.81 for Qwen3-8B. Metacognitive selection gives the highest cMFG* among the tested selection strategies for both models, although random selection has a lower BS for Llama3.1-8B-Instruct (0.23 versus 0.26).
In a 120-example study comparing sets of rewritten RLMF responses with FUT responses, three graduate-student annotators give the proposed method win rates of 98% for diversity, 98% for naturalness, 95% for helpfulness, and 96% for contextual suitability, counting ties at half weight. Reported Krippendorff’s alpha is 0.93; the preference task does not independently verify factual correctness.
6 Conclusion
The experiments show improved FC and self-assessment on the measured task, not broad metacognitive awareness. The Ethics Statement notes that faithful uncertainty and self-assessment do not replace verification of factual claims, particularly in consequential uses, and calls for safeguards and oversight.
Appendix
- Appendix A distinguishes the paper’s target from factual confidence calibration and surveys numerical and linguistic uncertainty methods. Appendix B details the LoRA-based pre-SFT and GRPO setups, all 10 datasets and their abbreviations, prompts, and judgment procedures: response-level faithfulness averages absolute differences between expressed and intrinsic sentence confidence, while cMFG* replaces potentially empty equal-width bins with width-weighted equal-mass bins over observed support.
- Appendix C specifies the GRPO objective and reward design, then tests reward formulations and weights, system prompts, advantage normalization, training-set size, and pre-SFT. Its RLMF ablations compare quadratic and linear self-assessment signals (cMFG* 0.84 versus 0.77 for Llama3.1-8B-Instruct), group-normalized signals, scaling the whole advantage or whole group, and using metacognition as an added reward; separate tests examine k and τ. Data-selection analyses compare high- and low-rated examples with sampled-response and active-learning variants.
- Appendix C also documents the human-rated score–hedge map and compares hedge counts, single-pass versus two-step editing, and rewriting models. Appendix D adds smaller-model results, per-bin faithfulness and training-trajectory analyses, and example generations; Appendix E describes the 120 human comparisons drawn from Natural Questions, SciQ, and SelfAware, including three responses per method, specified use contexts, and three annotators.
Brief Thoughts
Separating confidence training from editing permits different linguistic presentations of the same numerical scores without retraining the policy. The evidence remains specific to the paper’s operational measures: intrinsic confidence is sampled-response consistency, self-assessment concerns FC, and linguistic quality depends partly on a separate editor.