VideoAuto-R1: Video Auto Reasoning via Thinking Once, Answering Twice | Summary
08 Jan 2026 | Paper Review Video Question Answering Adaptive Computation Reinforcement LearningContents
- Summary
- 1 Introduction
- 2 Related Work
- 3 Preliminaries
- 4 VideoAuto-R1
- 5 Experiments
- 6 Conclusion
- Appendix
- Brief Thoughts
This article explains the key points of VideoAuto-R1: Video Auto Reasoning via Thinking Once, Answering Twice.
- 2026-01-08 (arXiv)
- Liu, Shuming, Zhuge, Mingchen, Zhao, Changsheng, Chen, Jun, Wu, Lemeng, Liu, Zechun, Zhu, Chenchen, Cai, Zhipeng, Zhou, Chong, Liu, Haozhe, et al.
- Meta AI, King Abdullah University of Science and Technology (KAUST), Princeton University
- Paper
- Github
- Project Page
Summary
- VideoAuto-R1: Video Auto Reasoning via Thinking Once, Answering Twice investigates when explicit chain-of-thought (CoT) benefits video understanding. Re-evaluation of three RL-trained video models shows that direct answering often matches or exceeds CoT on perception-oriented benchmarks, whereas VideoMMMU consistently benefits from reasoning.
- The method trains a single answer → think → answer sequence with verifiable rewards for both answers, greater weight on the reviewed answer, and a bonus for successful deferral. At inference, the model exits after its initial answer when its length-normalized token confidence meets a threshold; otherwise, it continues reasoning.
- With Qwen2.5-VL-7B, VideoAuto-R1 achieves 58.6% on VideoMMMU and reduces average response length from 149 tokens for standard RL with thinking to 44 tokens. Benefits are task-dependent: the reviewed-answer stage does not improve the reported grounding results, and some video QA scores remain below the corresponding base model.
1 Introduction
The introduction distinguishes symbolic reasoning from video perception: after relevant visual information is correctly extracted, many questions require little additional deduction. Unconditionally generating a long rationale adds decoding cost and can introduce errors rather than improve an answer.
- The proposed alternative separates learning to produce initial and reviewed answers from deciding whether to continue reasoning at inference.
- Figure 1 contrasts this design with always-thinking models and learned binary mode-switching approaches. The project and demo are available at https://ivul-kaust.github.io/projects/videoauto-r1.
The examples contrast a shirt-folding question that exits at confidence 0.99 with a pressure calculation that continues at 0.73 and changes C to D. They show how one output format supports both immediate answering and corrective reasoning.
2 Related Work
The related-work discussion covers CoT reasoning, video reasoning models, and adaptive thinking. VideoAuto-R1 optimizes both a direct answer and a reasoned answer during training rather than supervising a separate think/no-think decision.
2.1 Chain-of-Thought Reasoning
CoT prompting and reinforcement learning have improved mathematical, scientific, coding, and visual reasoning. The paper also cites evidence of overthinking in primarily perceptual tasks, motivating a test of whether these benefits extend to general video understanding.
2.2 Video Reasoning Models
Prior video reasoning methods extend R1-style training to QA, temporal grounding, object relations, and long-video narratives. Methods that use “thinking with frames” revisit selected visual content, whereas VideoAuto-R1 studies whether language-based reasoning should continue after an initial answer.
2.3 Auto-Thinking
Existing auto-thinking methods commonly learn switching policies through SFT or RL and may require balanced think/no-think data. VideoAuto-R1 avoids per-sample mode labels, specialized switch tokens or heads, and cold-start CoT SFT; selection uses the initial answer’s confidence at test time.
- The authors argue that scarce genuinely must-think video examples make supervised mode selection unstable. Their reproduced training-based baseline supports this concern for the tested setup, not for every possible learned routing approach.
3 Preliminaries
The preliminaries introduce GRPO training and compare direct and CoT inference in existing video reasoning models. The comparison motivates selective rather than unconditional reasoning.
3.1 Training Framework
Group Relative Policy Optimization (GRPO) samples candidate outputs, computes verifiable rewards, and normalizes them into group-relative advantages without a learned critic. A clipped importance-ratio objective updates the policy, with a KL penalty against a reference policy.
- The relative advantage is (A_i = \frac{r_i-\mu}{\sigma+\varepsilon}), where μ and σ are the reward mean and standard deviation within the sampled group.
- Standard rewards combine task correctness and output-format correctness. The tasks include QA, temporal grounding, and grounding QA.
The training corpus combines video QA and grounding examples with text and image math or science problems to support symbolic reasoning. After filtering, it contains approximately 83K samples; RL starts directly from the instruction-tuned backbone without a cold-start CoT SFT stage.
- Experiments with Video-R1-CoT SFT degraded the Qwen2.5-VL baseline and produced an inferior starting point for subsequent RL, supporting omission of that cold-start stage.
3.2 Analysis of Existing Video Reasoning Models
Table 1 re-evaluates Video-R1, Time-R1, and VideoChat-R1 with identical inputs: at most 256 frames and 16K total video tokens. Direct answering is competitive on VideoMME, LongVideoBench, and MMVU while producing far fewer tokens.
- For VideoChat-R1, CoT reduces VideoMME accuracy from 65.7 to 63.9 and LongVideoBench from 60.1 to 58.2, while increasing average response length from 4.3 to 126 tokens.
- All three models improve on VideoMMMU under CoT, by +1.0, +1.1, and +3.4 points respectively. Grounding is less uniform: CoT improves Time-R1 and VideoChat-R1 on Charades-STA but lowers Video-R1 from 42.0 to 34.9.
- The qualitative analysis associates useful reasoning with cleaner symbolic inputs and multi-step scientific deduction. Perception-oriented traces can repeat observations or hallucinate details.
Matched video inputs make this a within-model comparison of inference strategies. CoT consistently helps VideoMMMU, but lowers several perception-benchmark scores and substantially lengthens responses; grounding outcomes differ across models.
4 VideoAuto-R1
VideoAuto-R1 trains one sequence containing an initial answer, a rationale, and a reviewed answer. Figure 2 separates dual-answer GRPO training from the confidence-based inference decision.
The training branch rewards both answers within each sampled response before computing group-relative advantages. The inference branch makes a stopping decision after the initial answer rather than invoking a separately trained binary routing policy.
4.1 Thinking Once, Answering Twice
Each training response follows the strict format \\boxed{a1}<think>r</think>\\boxed{a2}: exactly two boxed answers and one intervening reasoning block, without additional text. Table 2 provides a prompt that elicits this sequence without cold-start SFT.
- When immediate answering is infeasible, the first box may contain the designated fallback text “Let’s analyze the problem step by step.” The inference rule forces continuation for this fallback instead of accepting it as an answer.
- The design separates how to reason, learned through RL, from when to reason, determined at inference. Users can also select the first or second answer unconditionally.
The prompt specifies the ordered output grammar and permits a fallback in the first box. It elicits both answer stages within one generation without adding a learned switching component.
4.2 Training: Dual-Answer Reward with GRPO
The dual-answer reward is R = w1 R_task^(1)(a1) + w2 R_task^(2)(a2) + λ R_fmt + α R_fallback, with w2 > w1 ≥ 0 and λ, α ≥ 0. Greater weight on the reviewed answer gives a higher reward to wrong → correct than to correct → wrong outcomes; it does not prevent harmful revisions.
- R_fallback is a binary bonus awarded when the initial answer is the designated fallback and the reviewed answer is correct. It rewards successful deferral rather than an incorrect initial guess.
- The format reward enforces the answer → think → answer grammar. Reviewed-answer task rewards generally exceed initial-answer task rewards during training.
4.3 Inference: Confidence-Based Early Exit
The initial answer’s confidence is s(a1) = (1/L) Σ_ℓ log pθ(tℓ | t<ℓ, q), computed over tokens inside its box. The model exits when s(a1) ≥ log τ; otherwise, it continues generating the rationale and reviewed answer.
- The fallback answer receives confidence −∞, ensuring that it cannot trigger early exit.
- This length-normalized log probability is not a calibrated probability of answer correctness. Initial answers usually contain fewer than ten tokens, limiting the work required to compute the score.
5 Experiments
The experiments evaluate accuracy, output length, and reasoning activation across video perception, video reasoning, temporal grounding, grounding QA, and image reasoning. Analyses test training strategy, adaptive routing, confidence thresholds, rewards, data curation, frame count, and cold-start SFT.
5.1 Experiment Details
The backbones are Qwen2.5-VL-7B-Instruct and Qwen3-VL-8B-Instruct. Training uses at most 4,096 video tokens and 256 frames, freezes the visual encoder, and updates the projector and LLM for one epoch on 32 H100 GPUs for approximately 35 hours.
- AdamW uses learning rate 1×10−6, weight decay 0.01, maximum gradient norm 1.0, and a constant schedule without warm-up. The global batch size is 256 and the KL coefficient β is 0.01.
- Reward coefficients are w1 = 0.9, w2 = 1.1, λ_fmt = 1, and α = 0.3. GRPO uses 16 rollouts per prompt at temperature 1.0.
- Testing uses lmms-eval, greedy decoding at temperature 0, a 4,096-token response limit, and τ = 0.97 tuned on a small held-out validation subset.
Video QA evaluation includes VideoMME without subtitles, MVBench, LongVideoBench, MMVU multi-choice, VideoMMMU, and MVP-mini. MVP reports pairwise accuracy, requiring correct answers for both visually similar videos with opposing answers.
- Charades-STA and ActivityNet measure temporal grounding through recall at IoU thresholds and mean IoU; NExT-GQA measures grounding QA.
- Qwen2.5-VL testing permits 16K video tokens with {64, 128, 256} frames; Qwen3-VL permits 128K tokens with {64, 256, 2048} frames. Main results report the highest performance across these settings, not one fixed frame configuration.
- Image evaluation uses MathVista, MathVision, MathVerse, MMMU, MMMU-Pro, and MM-Vet.
5.2 Main Results
Table 3 reports Qwen2.5-VL-based VideoAuto-R1 scores of 67.3 on VideoMME, 71.0 on MVBench, 60.5 on LongVideoBench, 69.7 on MMVU, 58.6 on VideoMMMU, and 39.4 on MVP. Reasoning-benchmark gains are larger than most perception-benchmark gains, but LongVideoBench decreases from the base model’s 60.9 to 60.5.
- The Qwen3-VL version reaches 65.0 on VideoMMMU and 43.0 on MVP, compared with base-model scores of 61.0 and 40.5. Its VideoMME and LongVideoBench scores, 71.7 and 67.4, remain below the baseline’s 72.5 and 67.6.
- Average response lengths are 44 and 52 tokens for the two backbones. For Qwen2.5-VL, reasoning activates on 25% of MVBench samples and 51% of VideoMMMU samples.
- The results support stronger reasoning-benchmark accuracy with selective decoding, not uniformly superior performance across all benchmarks. Table 3 reports 28% MMVU reasoning activation, whereas Tables 7–8 report 39%.
Both backbones improve on reasoning benchmarks while producing short average responses. LongVideoBench does not improve over either base model, and the Qwen3-VL version also falls below its baseline on VideoMME. The Qwen2.5-VL MMVU activation figure here is 28%, inconsistent with the 39% reported in Tables 7–8.
Table 4 shows grounding improvements over the Qwen2.5-VL baseline: Charades-STA mIoU rises from 52.9 to 60.0, ActivityNet from 26.9 to 47.6, and NExT-GQA accuracy from 53.3 to 80.6. Grounding evaluations use the initial answer by default because the reviewed answer does not improve the reported localization results.
- Qwen3-VL-based VideoAuto-R1 achieves 63.7 mIoU on Charades-STA, 51.9 on ActivityNet, and 81.1 accuracy with 44.2 mIoU on NExT-GQA.
- Comparisons with prior models are metric-specific: the Qwen2.5-VL version’s ActivityNet mIoU of 47.6 is below Time-R1’s 52.1, and its NExT-GQA mIoU of 36.7 is below VITAL’s 43.0.
Table 5 shows improvements over Qwen2.5-VL on all six evaluated image benchmarks under matched settings. MathVista increases from 69.4 to 73.7, MathVision from 26.3 to 29.6, MathVerse from 44.8 to 46.9, MMMU from 51.3 to 53.8, MMMU-Pro from 36.1 to 39.8, and MM-Vet from 60.0 to 61.9.
- The training corpus contains image reasoning data, so these results demonstrate applicability beyond video rather than transfer to a modality absent from training.
All six image-benchmark scores improve under matched evaluation settings. Because image reasoning examples are included in training, these results support multimodal applicability without isolating transfer to an unseen modality.
5.3 Analyses and Ablations
Table 6 compares SFT, RL without thinking, standard RL with thinking, and VideoAuto-R1 on the same training data. VideoAuto-R1 achieves 58.6 on VideoMMMU versus 56.4 for RL with thinking, while reducing average response length from 149 to 44 tokens.
- Relative to RL with thinking, VideoAuto-R1 improves VideoMME from 66.1 to 67.3, MMVU from 67.5 to 69.7, and Charades-STA from 59.8 to 60.0.
- This controlled comparison evaluates dual-answer training together with adaptive inference. It does not attribute the accuracy gain or the full token reduction to early exit alone.
With the same training data, VideoAuto-R1 improves the listed scores over standard RL with thinking and reduces response length from 149 to 44 tokens. The comparison evaluates the complete training-and-inference method, not early exit in isolation.
Table 7 compares inference-time confidence routing with a training-based baseline inspired by AdaptThink. The baseline labels examples using eight rollouts and maintains an approximately 1:1 think/no-think training balance, yet its auto mode activates reasoning on only 14% of evaluation samples.
- The training-based auto mode scores 55.7 on VideoMMMU and 36.8 on MVP, compared with VideoAuto-R1’s 58.6 and 39.4.
- Within VideoAuto-R1, always using the first answer requires 8 tokens on average, always using the reviewed answer requires 91, and adaptive inference requires 44.
- Adaptive inference nearly preserves reviewed-answer accuracy: VideoMMMU is 58.6 versus 58.7 for always-thinking, and MVP is 39.4 versus 39.8. This within-model comparison isolates the inference rule more directly than Table 6.
Table 8 relates initial-answer confidence to the benefit from reasoning. Mean initial-answer probability is 0.948 on MVBench, 0.933 on MMVU, and 0.874 on VideoMMMU; the table reports corresponding reasoning activation rates of 25%, 39%, and 51%.
- The reviewed answer improves accuracy over the initial answer by +0.1, +0.4, and +4.0 points respectively.
- Routing recall on cases where the initial answer is wrong but the reviewed answer is correct is 100% for MVBench and MMVU and 94% for VideoMMMU.
- High recall shows that routing captures most correctable cases in these evaluations. The table does not report routing precision or a full confidence-calibration analysis.
VideoMMMU combines lower initial confidence, more reasoning activation, and a larger reviewed-answer gain than MVBench or MMVU. High recall on wrong-first/correct-second cases indicates that routing captures most useful continuations, but does not establish routing precision or calibrated confidence.
Table 9 examines answer weights and the fallback bonus. With weights 0.9:1.1, adding the fallback reward improves VideoMMMU from 56.4 to 58.6, MVP from 37.2 to 39.4, and Charades-STA from 59.1 to 60.0.
- Asymmetric weighting alone is not uniformly better than 1:1: the 0.9:1.1 setting without fallback lowers VideoMME from 66.1 to 66.0 and MVP from 38.3 to 37.2.
- The best reported configuration combines modest reviewed-answer emphasis with the fallback bonus. Increasing weight asymmetry further does not consistently improve results.
The fallback bonus improves both tested asymmetric-weight configurations. Asymmetry without fallback yields mixed results, so the strongest reported configuration supports the combined reward design rather than asymmetry alone.
Figure 3 varies the early-exit threshold to show the accuracy–computation trade-off. Increasing τ from 0.86 to 0.98 raises VideoMMMU accuracy from 57.5% to 58.7% and reasoning activation from 29% to 55%.
- MVP accuracy rises from 39.16% to 39.37% while activation rises from 20% to 51%, a small accuracy return for the additional reasoning.
- VideoMME stays at 67.3% across the displayed thresholds even as activation increases. The paper adopts τ = 0.97 as a shared default.
Higher thresholds consistently increase reasoning activation. Accuracy improves on VideoMMMU, changes only slightly on MVP, and stays flat on VideoMME, showing task-dependent returns from additional decoding.
Figure 4 illustrates a corrective reasoning case involving a uniform distribution (X\sim U(1.5,4.5)). An initial answer D at confidence 0.92 is revised to C after calculating P(x < 3.5 | x < 4) = 0.8.
- The example demonstrates a successful correction on a symbolic question, but does not establish how frequently reasoning corrects answers or whether the generated rationales are generally faithful.
At confidence 0.92, the initial choice D is not accepted. The continuation computes the conditional probability under the uniform distribution as 0.8 and revises the answer to the ground-truth choice C.
6 Conclusion
The conclusion presents dual-answer training with confidence-based early exit as an alternative to unconditional video CoT. The experiments support reduced output generation with retained reasoning-benchmark accuracy, while showing little additional grounding benefit from the reviewed-answer stage.
Appendix
- A Training Data: The filtered dataset contains 6.4K text, 27.5K image, and 49.4K video examples. Sources include DAPO-Math, ViRL, ThinkLite-Hard, Video-R1, TVBench, STI-Bench, MMR-VBench, Charades-STA, ActivityNet, Time-R1, and NExT-GQA; the authors state that evaluation test samples are manually excluded.
- A Training Data: Filtering checks ground-truth validity, generates eight responses per example with Qwen2.5-VL-7B-Instruct, and uses Qwen3-30B-A3B-Instruct to judge correctness. QA samples with all eight responses correct or all eight incorrect are removed, while temporal grounding samples are retained. Under RL with CoT, filtering the multimodal corpus improves VideoMMMU from 55.4 to 56.4 and reduces the listed size from 138K to 83K; text-only training performs poorly on video tasks, and the full modality mixture performs best overall.
- B Reward Designs: QA rewards are binary, using math-verify for mathematical answers and normalized string comparison otherwise. Temporal grounding uses the maximum temporal IoU across predicted and reference segment pairs, with zero reward for an unparseable prediction; grounding QA sums answer correctness and temporal IoU, while strict regex checks enforce the output format. For binary QA rewards, weights 0.9 and 1.1 and fallback coefficient 0.3 give task-plus-fallback rewards of 0.9 for correct → wrong, 1.1 for wrong → correct, 1.4 for fallback → correct, and 2 for both answers correct, excluding the format term.
- C Prompt Template: The RL-without-thinking baseline produces only a boxed answer, whereas the RL-with-thinking baseline produces a <think>…</think> rationale followed by a boxed answer. These templates distinguish the baseline training strategies from VideoAuto-R1’s answer → think → answer format.
- D Training Curve: Initial- and reviewed-answer task rewards rise rapidly early in training and then improve more gradually for both backbones. Reviewed-answer rewards remain higher after convergence, with a larger separation for Qwen3-VL in the plotted curves. These observations do not isolate the causal effect of reward asymmetry or model capacity from other backbone differences.
- E Inference Strategy: Algorithm 1 greedily decodes until the first opening <think> tag, extracts the preceding boxed answer, and evaluates its mean token log probability. The implementation assigns the fallback a score of −1e6, accepts the initial answer when the score meets log τ, and otherwise resumes decoding from the existing prefix.
- F Additional Experiments — F.1 Performance with Different Frames: Increasing frames generally helps VideoMME and LongVideoBench more than VideoMMMU. For Qwen2.5-VL-based VideoAuto-R1, VideoMME increases from 64.6 at 64 frames to 67.3 at 256 frames, whereas VideoMMMU falls from 58.7 to 56.7. Table 15’s best VideoMMMU score of 58.7 differs slightly from Table 3’s reported 58.6.
- F.2 Analysis on Temporal Grounding Benchmarks: First-answer, second-answer, and adaptive inference produce identical listed ActivityNet and NExT-GQA results for Qwen2.5-VL-based VideoAuto-R1. The authors hypothesize that direct visual localization leaves little benefit for textual deduction and that omitting grounding-specific reasoning SFT limits refinement. These are proposed explanations, not separately tested mechanisms.
- F.3 Analysis of the Impact of Cold-Start SFT: SFT with Video-R1-CoT lowers VideoMME from 66.0 to 60.1, MVBench from 67.1 to 64.0, and VideoMMMU from 54.7 to 53.8. Subsequent RL from that checkpoint remains worse than direct RL with thinking, supporting omission of this particular cold-start dataset rather than a universal rejection of CoT SFT.
- G Limitations: Initial-answer confidence is not explicitly calibrated or optimized during training, and language-only CoT cannot revisit visual inputs to repair perception or temporal-boundary errors. The authors also identify limited benchmark scope and scarce genuinely must-think video data, restricting conclusions about long-range, causal, compositional, and counterfactual video reasoning.
- H Qualitative Examples: The appendix contrasts hallucinated CoT failures in perception tasks with a scientific QA case whose CoT answer matches the ground truth, unchanged grounding predictions, high-confidence early exits, and lower-confidence continuations. Grounding examples repeat intervals that differ from the annotations, demonstrating that agreement between initial and reviewed answers does not establish correct localization. The Belbin-role example confirms an initially correct answer rather than correcting it.
Brief Thoughts
The clearest contribution is the separation of dual-answer learning from inference-time allocation of reasoning. Table 7 provides direct evidence for the routing rule: adaptive decoding approaches always-thinking accuracy on VideoMMMU while reducing average output length from 91 to 44 tokens.
The efficiency result is a reduction in response generation, not a measured 3.3× end-to-end speedup: the reported full-method comparison is 149 versus 44 output tokens, and video encoding costs remain. Missing routing precision and calibration analysis, best-across-frame reporting, inconsistent activation figures, and mixed benchmark-level gains limit claims of universal superiority.