RbtAct: Rebuttal as Supervision for Actionable Review Feedback Generation | Summary
10 Mar 2026 | Paper Review Peer Review Generation Direct Preference Optimization Supervised Fine-TuningContents
- Summary
- 1 Introduction
- 2 Related Work
- 3 Review and Rebuttal Mapping
- 4 Methodology
- 5 Experiments
- 6 Conclusion
- Limitations
- Appendix
- Brief Thoughts
This article explains the key points of RbtAct: Rebuttal as Supervision for Actionable Review Feedback Generation.
- 2026-03-10 (arXiv)
- Wu, Sihong, Ma, Yiling, Zhao, Yilun, Hu, Tiansheng, Jiang, Owen, Patwardhan, Manasi, Cohan, Arman.
- Yale University, New York University, TCS Research
- Paper
- Github
Summary
- RBTACT targets perspective-conditioned segment-level review feedback generation: given a paper and a requested perspective, it produces one focused weakness or question. Its supervision comes from author rebuttals, distinguishing comments associated with reported revisions from those associated with plans, vague commitments, defenses, or deflections.
- The authors construct RMR-75K, containing 75,542 review–rebuttal mappings from 4,825 ICLR 2024 papers. They train Llama-3.1-8B-Instruct using 13,300 supervised examples followed by Direct Preference Optimization on 21,822 preference pairs, each restricted to the same paper and perspective.
- On ICLR 2025 evaluation data, RBTACT achieves the highest mean Actionability scores: 3.46/5 from human evaluators and 3.38/5 from GPT-5-chat. The reported paired human test supports an improvement over SFT-only, but not over GPT-5-chat; author uptake remains an imperfect proxy for feasible, scientifically valuable feedback.
1 Introduction
The paper addresses a recurring limitation of fluent AI-generated reviews: comments can lack concrete steps for improving a manuscript. RBTACT isolates individual weaknesses and questions and connects them to the author responses they elicited, rather than treating a mixed, full-length review as one supervision unit.
- Rebuttals supply implicit preferences based on author-reported uptake, rather than explicit human rankings of generated comments.
- Perspective conditioning separates concerns such as Experiments and Writing, supporting focused alignment and evaluation.
- The contributions are a segment-generation task, a labeled review–rebuttal dataset, and an SFT-to-DPO training pipeline.
2 Related Work
The related work separates peer-review datasets from feedback-generation methods. RBTACT uses author responses to train a generator, extending beyond collecting rebuttals or analyzing their discourse.
Peer Review Datasets.
PeerRead and NLPeer collect manuscripts and reviews, while ReviewMT, Review-5K, and DeepReview-13K support reviewer-model training. Rebuttal-aware resources include PRRCA, MOPRD, Re2, DISAPERE, and JitsuPeer; RMR-75K instead provides large-scale point-to-point alignment with perspective and impact labels for feedback generation.
- DISAPERE contains 506 review–rebuttal pairs with sentence-level discourse annotations, differing from RMR-75K in both annotation unit and scale.
- ARIES links comments to manuscript edits, whereas RBTACT links comments to rebuttal spans and derives training preferences from response categories.
Review Generation.
Prior review-generation systems use prompting, fine-tuning, structured pipelines, and multi-agent coordination to improve specificity. ReAct labels actionability and intent, ARIES studies implemented edits, and MARG generates actionable comments through specialized agents; RBTACT derives preferences from rebuttal-reported actions.
3 Review and Rebuttal Mapping
RMR-75K, expanded as Review-Map-Rebuttal, associates individual review concerns with the author-response spans that address them. Figure 1 summarizes collection, segmentation, alignment, filtering, and human verification—the preprocessing stages on which the downstream preference signal depends.
The pipeline makes alignment quality a prerequisite for training: public author responses are collected, critique units are isolated, spans are matched, and unreliable or non-substantive pairs are rejected. Unmatched concerns do not enter the retained supervision corpus.
3.1 Review–Rebuttal Source Data Collection
The source corpus consists of ICLR 2024 submissions collected through the official OpenReview API. Manuscript PDFs are converted to Markdown with MinerU; retained records include full paper text, complete reviews and free-form comments, and author rebuttal threads.
- Multiple author-response messages are concatenated chronologically into one response document per paper.
- The training corpus covers one conference year, not a cross-venue or cross-disciplinary sample.
3.2 Dataset Construction
Dataset Construction decomposes reviews into weakness/question segments and rebuttals into candidate response spans. The final alignment is one-to-one: each review segment and each rebuttal span can appear in at most one retained pair.
- Notation. Each mapped pair contains a review segment, an addressing rebuttal span, and a confidence score in [0, 1].
- Review Segmentation. Existing numbering or bullets supply the segments directly; otherwise GPT-5 splits Weaknesses and Questions into atomic critique units.
- Review–Rebuttal Mapping. Explicit references and quotations support a heuristic pass, followed by LLM-based semantic span matching. Greedy matching proceeds in descending confidence, and ties are discarded.
- Table 1 reports 75,542 mappings, 4,825 papers, 16,583 distinct reviews, 3.44 reviewers per paper, 15.66 mappings per paper, 4.56 mappings per review, and mean confidence 0.9268.
Multiple mapped concerns per review and paper provide candidates for within-paper preferences. The mean confidence of 0.9268 is an automatic matcher score, not independently measured alignment accuracy.
3.3 Cleaning and Quality Control
Cleaning and Quality Control retains substantive, confidently aligned concerns and removes unstable segmentations, unsupported matches, and responses that merely restate a question. The confidence threshold is denoted τ, but its numerical value is not reported.
- Structural Filter. Reviews are discarded when they lack visible itemization and cannot be segmented into stable atomic units.
- Coverage Filter. A segment is removed if it has no plausible rebuttal span with confidence ≥ τ.
- Confidence Filter. Retained pairs must meet the threshold; spans that repeat the question without addressing it are pruned.
- Substance Filter. Segments must raise a substantive issue or recommendation.
- Human Verification. Two annotators independently align 944 retained segments from 60 papers, with adjudication producing gold spans. Correctness requires token-level IoU ≥ 0.5.
- Table 2 reports Precision 0.93, Recall 0.90, F1 0.91, and pre-adjudication Cohen’s κ 0.80. These results validate alignment on the filtered sample, not coverage of concerns excluded by the pipeline.
Precision 0.93, Recall 0.90, and F1 0.91 support reliable matching within the filtered sample under token-level IoU ≥ 0.5. Cohen’s κ = 0.80 measures agreement between annotators before adjudication, not model–human agreement.
4 Methodology
RBTACT first learns perspective-specific segment generation through supervised fine-tuning, then uses DPO to favor comments associated with stronger author uptake. Rebuttals determine training labels and preferences; generation requires only the paper and requested perspective.
4.1 Task Definition
The task takes the complete paper and one target perspective and outputs a single focused weakness or question. The seven perspectives are Experiments, Writing, Presentation, Theory, Novelty, Reproducibility, and Evaluation.
- Review Perspective Labels. GPT-5 assigns labels automatically. Human spot checks report about 92% accuracy and Cohen’s κ = 0.81, with common confusions between Writing and Presentation and between Experiments and Evaluation.
- Rebuttal Impact Category as Actionability Signal. The categories are Concrete Revision Performed (CRP), Specific Revision Plan (SRP), Vague Commitment to Revise (VCR), Defend Without Change (DWC), and Deflect/Reframe (DRF).
- Impact labeling reports 89% accuracy against adjudicated human labels and inter-annotator κ = 0.79. These labels describe rebuttal content rather than independently verified manuscript changes.
- Figure 2 shows larger revision shares for Writing and Presentation and larger defense shares for Theory and Novelty.
Revision uptake depends on perspective: Presentation and Writing more often receive revisions, while Novelty and Theory more often receive defenses. This supports within-perspective pairing and cautions against interpreting revision frequency as a universal quality measure.
4.2 Training Dataset Construction
REVIEWSEG-SFT-13K contains 13,300 paper–perspective/segment examples from 4,637 papers, with exactly 1,900 instances per perspective. REVIEWPREF-DPO-22K contains 21,822 preference pairs from 4,825 papers, averaging 3,117 pairs per perspective.
- SFT Data. Targets are human reviewer segments selected from the mapped corpus; average paper and output lengths are 22,152 and 62 tokens.
- Preference Data. Comments are compared only within the same paper and perspective and only when their impact labels differ, following CRP > SRP > VCR > DWC > DRF.
- Pair counts are balanced across papers and perspectives, and reuse of individual segments is capped. Large, medium, and small impact gaps provide different preference difficulties.
- Table 3 also lists 1,330 extra SFT test pairs and 2,180 extra DPO test pairs, approximately 10% of the respective training-set sizes.
- Table 4 reports 10,268 large-gap pairs (47.1%) and 7,121 medium-gap pairs (32.6%). Its small-gap subtotal is printed as 6,433 (20.3%), although the listed rows sum to 4,433; the latter agrees with the overall total of 21,822.
The training examples combine long manuscript contexts with short outputs: average paper lengths exceed 21,000 tokens, while average comments contain 62–65 tokens. SFT is exactly balanced across perspectives; DPO reports average coverage rather than identical counts.
Large-gap comparisons dominate, particularly CRP versus DWC with 9,887 pairs. The printed small-gap subtotal of 6,433 conflicts with its rows, which sum to 4,433 and reconcile with the overall total of 21,822.
4.3 Policy Optimization
Direct Preference Optimization uses the frozen SFT checkpoint as its reference policy and learns from pairwise preferences without a separate reward model. The loss is L_DPO(θ) = −E[log σ(β(Δθ,ref(x, yw) − Δθ,ref(x, yℓ)))], where Δθ,ref(x, y) = log πθ(y | x) − log πref(y | x).
- The preferred response has the higher rebuttal-impact category; this ordering does not establish that it is the more severe or scientifically consequential critique.
- Stabilization. The objective adds positive-example negative log-likelihood: L(θ) = L_DPO(θ) + λE[−log πθ(yw | x)], with λ = 0.1, to limit drift in perspective control.
- The paper defines β > 0 as controlling preference sharpness but does not give its numerical setting in the reported training configuration.
4.4 Training Details
Both stages train Llama-3.1-8B-Instruct on NVIDIA H200 141GB GPUs with bf16 compute. SFT runs for 3 epochs, 4,989 steps, at learning rate 1.0×10−4 for approximately 120 hours; DPO runs for 2 epochs, 2,728 steps, at 1.0×10−5 for approximately 203 hours, with cosine scheduling in both stages.
- Appendix C specifies LLaMA-Factory and LoRA on q,k,v,o,gate,up,down projections, with ranks 8 for SFT and 16 for DPO.
- The context limit is 32,768 tokens and generation limit is 512 tokens. Table 9 defines effective batch sizes of 8 and 16 as per-device batch size multiplied by gradient accumulation.
- DPO combines FlashAttention-2, DeepSpeed ZeRO-2, gradient checkpointing, and a 4-bit quantized frozen reference model to accommodate long paper contexts.
5 Experiments
The experiments compare nine systems through human ratings, LLM pointwise ratings, LLM pairwise judgments, and reference-overlap metrics. They evaluate segment quality, not comprehensive review quality or observed improvements to revised manuscripts.
5.1 Baselines
The baselines are RBTACT-SFT; prompted GPT-5-chat, DeepSeek-V3.2, Llama-3.1-70B, and Qwen-3-32B; and task-adapted MARG, LimGen, and DeepReviewer-14B. Each receives paper text and a requested perspective, with its final output normalized to one review segment.
- Fine-tuned (SFT-only). RBTACT-SFT isolates the additional contribution of rebuttal-derived DPO.
- Prompted LLMs. The shared baseline prompt requests a weakness, question, or suggestion in 1 to 3 sentences.
- Other Methods. MARG retains the leader’s final synthesis, or its first bullet if necessary. LimGen retains the first limitation and its paired suggestion.
- DeepReviewer-14B runs without external tools; if it produces a full review, fixed heading-based extraction selects the matching subsection and its first paragraph.
- These adaptations standardize the output unit but evaluate restricted versions of systems designed for longer or multi-comment reviews.
5.2 Evaluation Protocol
The evaluation set contains 700 ICLR 2025 papers, with 100 paper–perspective instances per perspective and one human segment as each reference. The authors exclude resubmissions of ICLR 2024 papers to reduce overlap with the training source.
- Human Evaluation Setup. Three PhD-level or senior graduate annotators, each with ≥2 completed reviews at major ML venues, rate nine anonymized outputs on 50 papers. They use 1–5 scales for Actionability, Specificity, Groundedness, Relevance, and Helpfulness; scores are averaged across annotators.
- The human interface also asks which of two reviews is more useful and permits ties. The main text describes nine outputs as three randomized pairs, leaving the full presentation schedule unclear.
- LLM-as-a-Judge Evaluation Setup. GPT-5-chat evaluates 105 papers, yielding 945 pointwise evaluations across nine systems, and separately compares paired segments for Actionability.
- The LLM pairwise prompt forces a winner without ties, unlike the human interface. It requests a choice and justification rather than explicit per-candidate scores.
5.3 Experiment Results
Table 5 gives RBTACT the highest mean Actionability and Specificity in both studies: human scores are 3.46 and 4.08, and LLM scores are 3.38 and 3.70. Human and LLM Actionability are 3.28 and 3.18 for RBTACT-SFT and 3.38 and 3.28 for GPT-5-chat.
- Human Evaluation Results. GPT-5-chat scores higher on Groundedness, Relevance, and Helpfulness: 4.35, 4.98, and 4.47 versus RBTACT’s 4.30, 4.76, and 4.26.
- LLM-as-a-judge Pointwise Results. GPT-5-chat also leads those dimensions with 4.12, 4.95, and 3.78 versus RBTACT’s 4.05, 4.82, and 3.74.
- LLM-as-a-judge Pairwise Results. Table 6 reports RBTACT win rates of 57.1% against GPT-5-chat, 65.7% against RBTACT-SFT, and 76.2% against LimGen.
- Figure 17 shows perspective-dependent results: RBTACT wins 40.0% of Experiments comparisons against GPT-5-chat, despite its overall advantage. Several baseline-to-baseline entries in the overall heatmap differ from Table 6.
- Human–LLM Agreement. Model-level Actionability rankings have Spearman’s ρ = 0.94 and Kendall’s τb = 0.87; matched item scores across 20 papers × 9 models have Spearman’s ρ = 0.52. High ranking agreement therefore coexists with only moderate item-level agreement.
- Appendix F.3 reports an Actionability advantage over RBTACT-SFT with p = 0.018 and rank-biserial effect 0.23. Differences from GPT-5-chat (p = 0.276) and DeepReviewer-14B (p = 0.086) are not significant in the reported tests; no multiple-comparison adjustment is described.
RBTACT has the highest Actionability and Specificity means in both evaluation modes, while GPT-5-chat leads Groundedness, Relevance, and Helpfulness. The result is dimension-specific rather than overall dominance.
RBTACT wins more than half of the forced-choice Actionability comparisons against each baseline, including 57.1% against GPT-5-chat. These are decisions from one LLM judge and should not be interpreted as significant human preference differences.
The reported paired tests support an Actionability improvement over SFT, but not over GPT-5-chat or DeepReviewer-14B. They also report lower Helpfulness than GPT-5-chat; the PDF does not describe a multiple-comparison adjustment.
Aggregate advantages are not uniform across perspectives: RBTACT’s win rate against GPT-5-chat is 40.0% for Experiments and 80.0% for Theory. Several baseline-to-baseline entries in the overall panel differ from Table 6, so those two presentations are not fully reconciled.
5.4 Automatic Evaluation
Automatic Evaluation compares generated comments with one human reference on all 700 instances using BLEU@4, ROUGE-Lsum (F1), METEOR, and chrF. Table 10 gives RBTACT the highest ROUGE-Lsum, 12.64, and METEOR, 11.65; RBTACT-SFT leads BLEU@4 at 14.93, while GPT-5-chat leads chrF at 24.90.
- RBTACT scores 14.62 on BLEU@4 and 18.57 on chrF; all table values are reported as percentages.
- The small SFT-to-DPO overlap changes do not directly measure actionability. A valid critique can differ substantially from the single reference comment.
DPO slightly raises ROUGE-Lsum and METEOR relative to SFT while slightly lowering BLEU@4. GPT-5-chat leads chrF, showing that reference-similarity rankings depend on the metric and do not directly measure implementability.
5.5 Case Study
Figure 18 illustrates the output style encouraged by training through Experiments, Presentation, and Evaluation examples. RBTACT names manipulations, reporting locations, metrics, and procedures, while the comparison comments provide broader recommendations.
- The Experiments example requests retraining without MixUp/CutMix under a fixed seed with three independent trials, mean±std Top-1 in Table 3, and a Corrupted ImageNet check.
- The Presentation example identifies Figs. 2–3 and requests larger labels, explicit y-axis units, legends below panels, captions defining metrics and sample sizes, and an OKLCH-based color-blind-safe palette.
- The Evaluation example specifies Sec. 4.2, identical prompts for baselines and SOTA, macro-F1 and calibration, 95% CIs via paired bootstrap over papers, and an Appendix error taxonomy.
- The selected cases demonstrate explicit instructions, not independent verification of their feasibility or grounding. The fixed-seed wording in the first example also leaves the source of variation across trials unspecified.
The examples show the training target through named experimental factors, reporting locations, figure edits, metrics, and confidence-interval procedures. Their explicitness illustrates Actionability but does not verify the correctness or feasibility of the instructions.
5.6 Summary
The evidence supports a targeted Actionability improvement from rebuttal-derived training at 8B scale, with competitive Specificity. Claims of broader parity require qualification: GPT-5-chat has higher pointwise Groundedness, Relevance, and Helpfulness, and its human Helpfulness advantage is supported by the reported paired test.
- Appendix F.1 uses perspective as a severity proxy: major-topic mappings have DWC 45.4%, compared with 20.9% for minor topics, but still show CRP+SRP uptake of 50.7%. This demonstrates coverage of substantive topics, not directly measured issue severity.
- Paper-strength tertiles each contain 35 papers. Table 12 reports Actionability gains of +0.27 for Weak, +0.22 for Medium, and +0.19 for Strong papers, but does not identify the baseline system.
- A same-perspective nearest-neighbor retrieval baseline scores 3.02 on LLM Actionability versus RBTACT’s 3.38 and trails on all five dimensions. This supports learned generation over that particular direct-reuse baseline.
Each 35-paper strength bucket has a positive reported Actionability difference, largest for weaker papers. Because the comparator is labeled only Baseline, the table cannot establish which system RBTACT exceeds within each bucket.
Same-perspective nearest-neighbor retrieval trails RBTACT on all five rated dimensions, including Actionability 3.02 versus 3.38. This supports adaptation through generation over direct reuse; RBTACT’s Specificity is printed as 3.71 here rather than Table 5’s 3.70.
6 Conclusion
The conclusion presents rebuttals as a supervision source for focused feedback and announces the release of RMR-75K and code under the RbtAct identifier. The experiments support improvements in rated segment quality, but do not show that generated comments cause better revisions or improve final scientific validity.
Limitations
The paper acknowledges that rebuttals measure short-horizon author uptake rather than long-term implementation and may include strategic promises or deferrals. Generalization beyond public OpenReview rebuttals, computer science, and English remains uncertain, and the current setup lacks rigorous verification against manuscript, code, or data artifacts.
- An author’s defense can be justified, while an accepted revision can address an easy or low-value concern. The impact ordering is therefore a noisy preference assumption, not a direct measure of scientific merit.
- Filtering to answered, confidently aligned segments can exclude difficult or ignored concerns; the alignment validation does not assess those exclusions.
- Precise instructions can still be infeasible. Actual execution of generated revision plans is not evaluated.
Appendix
- A Review and Rebuttal Mapping Details supplies two mappings from “Large Language Models as Tool Makers” and the preprocessing prompts. A.1 Example of Mapping illustrates extracted response spans; A.2 Prompts for Review Segment and Mapping documents segmentation and verbatim extraction. The mapping prompt permits copying one response for multiple concerns, whereas the main method imposes one-to-one final alignment, so post-processing is necessary to enforce the stated constraint.
- B Label Details provides definitions in B.1 Detailed Definition of Perspective labels and Impact Categories and counts in B.2 Perspective and Impact Category Distribution: Experiments is largest with 25,160 mappings; overall counts are CRP 33,502 / 44.3%, SRP 7,574 / 10.0%, VCR 2,046 / 2.7%, DWC 31,047 / 41.1%, and DRF 1,373 / 1.8%. B.3 Prompt for labeling perspectives includes Miscellaneous beyond the seven generation perspectives and requests only a category name, unlike the main text’s description of a short rationale. B.4 Prompt for labeling impact categories distinguishes completed revisions or supplied artifacts from plans, vague promises, defenses, and deflections.
- C Additional Training Details covers Framework and hardware., Memory and speed optimizations., and Schedules., documenting LLaMA-Factory, H200 GPUs, LoRA adapters, and the long-context DPO configuration. Table 9 specifies LoRA α = 16 and dropout = 0.05 in both stages, warmup ratios 0.10 and 0.05, and per-device batch size 1 with accumulation 8 and 16 for SFT and DPO. The reported training is parameter-efficient adapter tuning rather than full-parameter fine-tuning.
- D Evaluation Details supplies D.1 Baseline Prompt and D.2 Baseline Adaptation Details for MARG., LimGen., and DeepReviewer-14B., including deterministic extraction when outputs contain multiple comments. D.3 Human Expert Evaluation Protocol separates pairwise usefulness preferences from per-candidate ratings in D.3.1 Interface; D.3.2 Dimension Rubrics (Concise) distinguishes implementable steps, precise pointers, evidence support, perspective alignment, and constructive utility. D.4 LLM-as-a-Judge Evaluation documents pointwise scoring and forced-choice Actionability comparisons, while D.5 Automatic Evaluation reports overlap metrics; GPT-5-chat serves as both a competing generator and the sole reported LLM judge.
- E Case Study presents the three comparisons in Figure 18. The examples illustrate how explicit factors, locations, metrics, confidence-interval procedures, and deliverables turn broad recommendations into instructions, but they are selected illustrations rather than a representative estimate of improvement.
- F Additional Analyses examines perspective-based severity and OpenReview-rating tertiles in F.1 Severity and Paper-Strength Analyses.; F.1.1 groups Experiments, Evaluation, Theory, Novelty, and Reproducibility as major topics, while F.1.2 Setup. and Interpretation. report positive gains across strength buckets without identifying the comparator. F.2 Retrieval Baseline Analysis encodes the test paper, searches same-perspective training comments while excluding the same paper, and returns the nearest comment without rebuttal access. F.3 Ordinal-Aware Analysis of Human Ratings reports medians, IQRs, rates ≥4, pooled-distribution ridit scores, and paired Wilcoxon tests: RBTACT’s Actionability median is 4 versus 3 for all baselines, with 52% of ratings ≥4; its Helpfulness difference from SFT is not significant (p = 0.641), while GPT-5-chat is significantly better (p = 0.031, effect −0.19), and some Table 14 means differ from Table 5.
Brief Thoughts
The useful methodological contribution is recovering preferences from reviewer–author interactions while keeping comparisons within a paper and perspective. The SFT ablation and paired human test support an additional Actionability benefit from this signal, but not superiority over GPT-5-chat or selection of the most scientifically important concerns. Author willingness to revise is distinct from critique validity and value. The perspective-based severity analysis addresses topic coverage but does not resolve that distinction. Observed manuscript changes and artifact-level verification would test the intended outcome more directly than reference overlap or selected examples. Unreported numerical settings and internal table discrepancies limit exact reproduction despite substantial implementation detail.