Gorio Tech Blog search

PhyGDPO: Physics-Aware Groupwise Direct Preference Optimization for Physically Consistent Text-to-Video Generation | Summary

|

Contents

This article explains the key points of PhyGDPO: Physics-Aware Groupwise Direct Preference Optimization for Physically Consistent Text-to-Video Generation.

  • 2025-12-31 (arXiv)
  • Cai, Yuanhao, Li, Kunpeng, Jia, Menglin, Wang, Jialiang, Sun, Junzhe, Liang, Feng, Chen, Weifeng, Juefei-Xu, Felix, Wang, Chu, Thabet, Ali, et al.
  • Meta Superintelligence Labs, Johns Hopkins University, Meta BizAI, CUHK
  • Paper
  • Github
  • Project Page

Read this article in Korean


Summary

  • PhyGDPO: Physics-Aware Groupwise Direct Preference Optimization for Physically Consistent Text-to-Video Generation studies post-training for physical plausibility without extending prompts at inference. It combines physics-rich real-video selection with a preference objective that treats real videos as winners and generated videos as losers.
  • PhyAugPipe filters one million text-video pairs into PhyVidGen-135K and resamples 17K pairs for training. PhyGDPO derives a tractable surrogate from a groupwise Plackett–Luce objective, uses Physics-Guided Rewarding (PGR) to emphasize difficult samples, and shares a frozen backbone between policy and reference through LoRA-Switch Reference (LoRA-SR).
  • Post-training Wan2.1-14B raises the VideoPhy2 overall score from 0.1288 to 0.1627 and the hard-action score from 0.0111 to 0.0500; PhyGenBench average performance reaches 0.55 versus VideoDPO’s 0.54. Human preferences and cross-evaluator tests support improved perceived physical plausibility, but neither the objective nor these evaluations establish guaranteed physical correctness.

1 Introduction

The introduction distinguishes visual quality from physically consistent motion and interaction. It argues that graphics-based supervision is difficult to extend to complex real scenes, prompt extension depends on fallible external reasoning, and supervised training lacks explicit negative examples that discourage implausible generations.

  • The proposed approach addresses scarce physics-rich preference data, unreliable generated winners and isolated pairwise comparisons, and the memory cost of duplicating a full reference model.
  • Figure 1 motivates the task with selected gymnastics, soccer, basketball, and glass-smashing examples. These illustrate visible failure modes rather than establish aggregate superiority over closed-source models.

The selected frame pairs expose concrete failure modes in body structure, limb–object contact, ball–hoop interaction, and glass breakage. They motivate physics-oriented post-training, but the examples alone do not establish comparative success rates.

Selected gymnastics, soccer, basketball, and glass-smashing generations comparing Wan2.1-14B, Wan2.1-14B + PhyGDPO, and closed-source alternatives.
Selected gymnastics, soccer, basketball, and glass-smashing generations comparing Wan2.1-14B, Wan2.1-14B + PhyGDPO, and closed-source alternatives.

2 Method

The method couples data construction with preference post-training. PhyAugPipe selects real videos containing informative physical interactions, while PhyGDPO contrasts each selected video with generated alternatives under the same original prompt.

  • VLMs supply filtering and difficulty scores for data preparation and preference weighting; inference does not require VLM-generated physics explanations.
  • Real footage determines the winning label. VLM scores control sample selection and weighting rather than choose the winner.

2.1 Physics-Augmented Video Data Construction

Data Filtering with CoT uses Qwen2.5-72B-Instruct to parse entities, actions, forces, and outcomes, check them against video frames, reason about causal interactions, and assign physics_richness in [0, 1]. Algorithm 1 sets physics_label = 1 when physics_richness ≥ 0.60 and limits the extended prompt to at most 100 words.

  • The scoring rules reward explicit interactions and causal clarity while penalizing camera motion, stylization, and static aftermath.
  • Extended prompts are retained for other research uses, but PhyGDPO uses original human-written prompts to avoid introducing rewriting errors or bias.

Action Clustering via Semantics Matching assigns filtered prompts to a predefined list of challenging action categories using a sentence Transformer. Data Sampling with Physics Rewarding allocates more real-video examples to categories on which the base generator performs poorly, as summarized in Figure 2.

  • For each category, top-nc semantically matched representatives provide prompts for assessing generation difficulty with VideoCon-Physics. Semantics adherence and physics commonsense scores are averaged within each generated video and then within each category to obtain Sf(k).
  • The sampling rule uses dk = 1 − Sf(k), rk = exp(τ dk), and Hr(k) = min(Hf(k), N rk / Σj rj). The cap prevents a category from contributing more examples than it contains.

The pipeline separates physics-richness filtering from difficulty-aware allocation. It rejects low-information scenes, groups retained prompts by action semantics, and shifts the training distribution toward actions for which the generator receives low physical and semantic scores.

PhyAugPipe stages: source data, CoT analysis, physics filtering, action clustering, and physics-rewarded sampling.
PhyAugPipe stages: source data, CoT analysis, physics filtering, action clustering, and physics-rewarded sampling.

2.2 Physics-Aware Groupwise Direct Preference Optimization

Groupwise Probability assigns a real winner a softmax probability against m generated losers using the Plackett–Luce model. Substituting the DPO reward parameterization and cancelling the prompt-dependent partition term gives a groupwise loss based on policy-to-reference log-likelihood ratios.

  • With fθ(x, c) = log(pθ(x|c)/pψ(x|c)), Eq. (4) is LGDPO = E[log(1 + Σj exp(β(fθ(x0lj, c) − fθ(x0w, c))))]. A generated alternative with a large relative advantage increases the penalty.
  • This formulation models the winner’s probability over a candidate group; it does not require a complete ranking among the generated losers.

The paper obtains an efficient surrogate in two stages: Jensen’s inequality bounds the multi-timestep expression by a single-timestep loss, and a second inequality converts the group expression into a sum of weighted pairwise terms. Each iteration can therefore sample one winner-loser pair and one timestep instead of evaluating every group member.

  • The intermediate single-timestep group objective still requires 2m + 2 model evaluations, counting policy and reference evaluations for the winner and m losers.
  • The key inequality is 1 + Σj exp(xj) ≤ Πj(1 + exp(αj xj))γj, under 0 < αj ≤ 1 and γj ≥ 1/αj. Eq. (8) expresses the resulting upper bound as an expectation of −mγj log σ(αjβT(Δkw − Δklj)); the factor m corresponds to uniform sampling of the loser index j.
  • Training optimizes an upper-bound surrogate, not the original groupwise likelihood directly. The paper does not report the tightness of the bound.

Physics-Guided Rewarding maps each generated loser’s normalized semantics adherence and physics commonsense scores to difficulty vj = 1 − (sjsa + sjpc)/2. Sigmoid functions increase comparison sharpness αj and loss weight γj for lower-scoring samples while keeping both coefficients bounded.

  • Eq. (9) uses αj = αmin + (1 − αmin)σ(κα(vj − bα)) and γj = [1 + λσ(κγ(vj − bγ))]/αmin.
  • With 0 < αmin ≤ 1 and λ ≥ 0, these definitions satisfy the inequality’s coefficient constraints because αj ≥ αmin and γj ≥ 1/αmin ≥ 1/αj. The reported settings satisfy these conditions.
  • Difficulty combines prompt adherence with physical plausibility, so PGR is not a physics-only reward. Increasing its coefficients does not by itself guarantee a larger gradient for every sample.

Flow Matching Probabilistic Modeling uses rectified-flow interpolation xt = (1 − t)x0 + tx1, with x1 drawn from N(0, I) and target velocity x1 − x0. A vanishing-noise Gaussian approximation converts policy-reference transition log-ratios into differences between their squared velocity-prediction errors.

  • The final Eq. (13) objective is L = E[−mγj log σ(−αjβT[(ℓθw − ℓψw) − (ℓθlj − ℓψlj)])]. It favors a greater policy-relative error improvement on the real winner than on the generated loser.
  • The scale h²/(2ε) from the Gaussian approximation is absorbed into β. This supplies an approximate probabilistic treatment of deterministic rectified-flow transitions.
  • Eq. (5) identifies a terminal-video log-likelihood ratio with a sum of transition log-ratios. Such a sum directly describes a trajectory likelihood ratio; the paper does not derive the marginalization needed to identify it with the terminal-video ratio.

LoRA-Switch Reference freezes the backbone and attaches trainable LoRA modules to the query, key, value, and output projections of each self-attention layer. An environment manager disables the adapters for reference evaluation and enables them for policy evaluation, avoiding a second full backbone.

  • Eq. (14) uses Y = X(W + 1action(α/r)BA)⊤, where 1action = 0 selects the reference and 1action = 1 selects the trainable policy.
  • Figure 3 connects the shared-backbone mechanism to real-video winners, multiple seeded generated losers, and VLM-based sample weighting.
  • The design reduces both trainable parameters and reference duplication. Its performance ablation therefore tests the combined adaptation and reference-sharing configuration.

The diagram distinguishes two roles: real footage fixes the winning label, while VLM scores adjust the weighting of generated alternatives. Policy and reference evaluations share the frozen backbone, with LoRA modules enabled only for the policy.

PhyGDPO architecture combining LoRA-Switch Reference, real-video winners, generated losers, and physics-guided groupwise preference training.
PhyGDPO architecture combining LoRA-Switch Reference, real-video winners, generated losers, and physics-guided groupwise preference training.

3 Experiment

The experiments filter one million high-quality text-video pairs into 135K physics-rich pairs, then use action clustering and rewarded resampling to obtain 17K training pairs. Wan2.1-14B generates alternatives from these 17K prompts with different random seeds, and VideoCon-Physics scores the generated videos.

  • The sampling budget is N = 20000, whereas the resulting resampled dataset contains 17K pairs; the per-category rule caps sampling by available data.
  • Evaluation uses short, unextended prompts from VideoPhy2 and PhyGenBench. The paper does not specify the numerical loser-group size m.

VideoPhy2-AutoRater evaluates semantics adherence and physics commonsense on scales from 1 to 5; the reported score is the fraction of videos with both scores ≥ 4. PhyGenBench evaluates the occurrence, order, and naturalness of required phenomena across 27 physical laws in 4 domains using GPT-4o, CLIP, InternVideo2, and VideoLLaVA.

  • VideoPhy2-AutoRater differs from the VideoCon-Physics model used to score training examples, reducing direct dependence on the training scorer.
  • These benchmarks assess physical plausibility through model judgments rather than measure physical quantities or conservation-law residuals.

Wan2.1-14B is post-trained in PyTorch for 10K steps with batch size 8 on 8 H100 GPUs for 6 days, using BF16 mixed precision and sublinear memory training. Training and inference resolution is 480×832.

  • AdamW uses β1 = 0.9, β2 = 0.999, and weight decay 0.01; cosine annealing lowers the learning rate from 1e−5 to 1e−6.
  • Reported settings are τ = 3, N = 20000, αmin = 0.5, kγ = 2.0, bγ = 0.4, λ = 0.6, kα = 5.0, and bα = 0.5. LoRA rank and scale factor are both 48.

3.1 Comparison with State-of-the-Art Methods

Table 1 reports the highest VideoPhy2 overall score for PhyGDPO, 0.1627, compared with Wan2.1-14B at 0.1288, VideoDPO at 0.1373, PhyT2V at 0.1492, Sora2 at 0.1508, and Veo3.1 at 0.1525. Its hard-action score is 0.0500, but it does not lead every track: Veo3.1 scores 0.1887 on Interaction versus PhyGDPO’s 0.1761.

  • The hard-action score is approximately 4.5× the base model’s 0.0111, though the absolute fraction meeting both evaluation thresholds remains low.
  • On PhyGenBench, PhyGDPO reaches Avg = 0.55 versus VideoDPO’s 0.54 and PhyT2V’s 0.50. It leads Mechanics at 0.55 and Thermal at 0.58, ties VideoDPO on Optics at 0.60, and trails VideoDPO on Material, 0.47 versus 0.58.
  • The tables support overall gains, but the PhyGenBench average advantage over VideoDPO is small and no uncertainty estimates accompany these automated benchmark scores.

The user study includes 104 participants, each completing 48 trials with 24 prompts from VideoPhy2 and 24 from PhyGenBench. Participants choose the video that better follows physical laws; same-prompt comparisons are anonymous and randomly ordered, and the paper reports 95% confidence intervals no wider than ±2.4%.

  • Table 1 reports preferences for PhyGDPO of 94.2% over Vcrafter2, 89.4% over VideoDPO, 88.5% over PhyT2V, 86.5% over Wan2.1-14B, 82.7% over Hunyuan, 67.3% over Sora2, and 64.4% over Veo3.1.
  • Human judgments corroborate perceived physical realism beyond agreement with one VLM evaluator. They do not isolate physical reasoning from visual coherence or other perceptual cues.

The qualitative comparisons cover human actions, object interactions, and physical phenomena. Figures 4 and 5 show selected gymnastics, polo, squash, and handspring sequences; Figure 6 extends the comparison to water refraction and paper combustion.

  • The examples emphasize body-shape continuity, contact between tools or limbs and objects, and plausible event progression rather than image quality alone.
  • The displayed frames illustrate these behaviors but cannot fully establish temporal dynamics. Video results are linked at https://caiyuanhao1998.github.io/project/PhyGDPO.

Gymnastics tests body-shape continuity during a difficult maneuver, while polo tests coordination among rider, horse, mallet, and ball. The selected PhyGDPO frames illustrate the intended interaction improvements without quantifying their frequency.

VideoPhy2 qualitative comparisons for gymnastics and polo across six methods.
VideoPhy2 qualitative comparisons for gymnastics and polo across six methods.

The squash example focuses on racket–ball contact, and the handspring example focuses on coordinated body movement around a platform. These user-input examples broaden the qualitative evidence beyond benchmark prompts, but do not constitute a systematic generalization test.

Qualitative comparisons for user-input squash and handspring prompts.
Qualitative comparisons for user-input squash and handspring prompts.

PhyGDPO leads the reported overall benchmark scores and is preferred in each listed human comparison. The domain breakdown limits the claim: Veo3.1 leads VideoPhy2 Interaction, VideoDPO leads PhyGenBench Material, and the PhyGenBench average advantage over VideoDPO is only 0.01.

VideoPhy2 and PhyGenBench quantitative results with pairwise human preferences for PhyGDPO.
VideoPhy2 and PhyGenBench quantitative results with pairwise human preferences for PhyGDPO.

The examples test optical appearance changes when a pencil enters water and event progression when burning paper reaches dry paper. They illustrate the authors’ claimed refraction and flame-propagation improvements, but contain no quantitative optical or combustion measurements.

PhyGenBench comparisons for water refraction and paper combustion.
PhyGenBench comparisons for water refraction and paper combustion.

3.2 Ablation Study

With the number of selected training examples held fixed, Table 2a progressively adds CoT filtering, action clustering, and rewarded sampling to a random-data baseline trained with PhyGDPO. VideoPhy2 overall scores rise from 0.1475 to 0.1525, 0.1575, and 0.1627; hard-action scores rise from 0.0333 to 0.0500.

  • This comparison supports the utility of data selection and allocation beyond dataset size alone.
  • Because the ablation is cumulative, it measures each component’s contribution in the preceding configuration rather than all possible component interactions.

Table 2b progressively adds LoRA-SR, groupwise modeling, and physics rewarding to Wan2.1-14B, increasing the overall score from 0.1288 to 0.1458, 0.1559, and 0.1627. Figure 7 illustrates the progression with a tennis ball that floats more stably rather than sinking or bouncing unrealistically.

  • Under the same data and training settings on Wan2.1-1.3B, Table 2c reports overall scores of 0.1283 for Flow-DPO, 0.1305 for VideoDPO, and 0.1407 for PhyGDPO.
  • Hard-action scores are 0.0296, 0.0278, and 0.0444, respectively. Figure 8 shows the corresponding selected weightlifting comparison for a prompt specifying a 25kg barbell.

Table 2d reports that LoRA-SR changes GPU memory from 48.7GB to 25.3GB and stored weights from 5.3GB to 84MB on Wan2.1-1.3B, while improving the overall score from 0.1373 to 0.1407. Figure 9 illustrates improved hand–shovel interaction in a snow-shoveling example.

  • LoRA-SFT uses 24.7GB and 84MB but reaches only 0.1283 overall and 0.0167 on hard actions, compared with PhyGDPO’s 0.1407 and 0.0444.
  • The prose claims a 44% memory reduction, which does not match the listed 48.7GB and 25.3GB values. The absolute table values are the clearer evidence.
  • The with/without LoRA-SR comparison changes the update parameterization as well as reference storage, so the performance difference cannot be attributed solely to switching reference modes.

Cross-evaluation preserves the method’s overall advantage when changing either the evaluator or the generator. Table 2e reports a Gemini-2.5-pro overall score of 0.5525 for PhyGDPO, versus 0.5373 for Sora2 and 0.5034 for Veo3.1; Table 2f reports 0.1426 for Vcrafter2 + PhyGDPO, versus 0.1373 for Vcrafter2 + VideoDPO.

  • The alternative evaluator supports an improvement not specific to VideoPhy2-AutoRater, although both evaluations rely on VLM judgments.
  • The matched-setting Vcrafter2 experiment supports applicability beyond the main Wan2.1 backbone. Table 2 collects these cross-evaluations alongside the data, objective, and efficiency ablations.

The six panels evaluate data selection, cumulative objective changes, matched-setting DPO comparisons, adapter efficiency, alternative scoring, and another backbone. The clearest efficiency evidence is the absolute change from 48.7GB to 25.3GB GPU memory and from 5.3GB to 84MB storage.

Component ablations, controlled DPO comparisons, LoRA-SR resource measurements, and cross-evaluations on VideoPhy2.
Component ablations, controlled DPO comparisons, LoRA-SR resource measurements, and cross-evaluations on VideoPhy2.

The tennis-ball sequence illustrates the cumulative ablation: the complete method shows more stable surface floating than configurations with sinking or excessive bouncing. It depicts buoyancy-related behavior rather than measuring buoyant forces.

Progressive PhyGDPO component additions for a tennis ball placed on water.
Progressive PhyGDPO component additions for a tennis ball placed on water.

Table 3 varies one PGR hyperparameter at a time while fixing the others at their selected values. The best reported setting is αmin = 0.5, kα = 5.0, bα = 0.5, λ = 0.6, kγ = 2.0, and bγ = 0.4, yielding 0.1627 overall.

  • The tested overall scores range from 0.1579 to 0.1627, all above the 0.1575 configuration without rewarded data sampling in Table 2a. That comparison does not isolate optimization-level PGR, whose cumulative ablation appears in Table 2b.
  • The table provides local sensitivity evidence, not a joint hyperparameter search or an uncertainty analysis of the small score differences.

Each panel varies one rewarding parameter around the selected configuration. All displayed scores exceed 0.1575, but the narrow differences have no accompanying uncertainty estimates, so the table supports local sensitivity assessment rather than precise parameter optimality.

One-at-a-time sensitivity analysis of the six PGR hyperparameters on VideoPhy2.
One-at-a-time sensitivity analysis of the six PGR hyperparameters on VideoPhy2.

Using the same Wan2.1-T2V-1.3B backbone, training data, and settings makes this example more controlled than the closed-source comparisons. The PhyGDPO sequence shows a more coherent overhead lift, although the specified 25kg load is not physically verified by the visual evaluation.

Matched-setting baseline, VideoDPO, Flow-DPO, and PhyGDPO generations for a snatch with a 25kg barbell.
Matched-setting baseline, VideoDPO, Flow-DPO, and PhyGDPO generations for a snatch with a 25kg barbell.

The with/without comparison highlights hand shape and contact with the shovel during snow removal. It illustrates the qualitative benefit of the LoRA-SR training configuration, but cannot separate reference sharing from the effect of restricting updates to adapters.

Snow-shoveling generations with and without LoRA-SR on Wan2.1-T2V-1.3B.
Snow-shoveling generations with and without LoRA-SR on Wan2.1-T2V-1.3B.

The related-work section covers text-to-video modeling and preference optimization. It positions PhyGDPO as physical-plausibility post-training using real-video preference supervision, without inference-time prompt extension.

4.1 Text-to-Video Generation Models

The paper reviews Transformer-based video generators that predict diffusion noise or flow-matching velocity fields. It contrasts PhyGDPO with DiffPhy and PhyT2V, which use LLM-generated descriptions of physical laws and phenomena for iterative generation or fine-tuning.

  • PhyGDPO still uses VLM reasoning during data preparation, but it retains original prompts for training and does not require prompt extension at inference.
  • The experiments measure generated outputs rather than directly test whether the model has acquired an internal physics-reasoning process.

4.2 Direct Preference Optimization

The DPO discussion traces preference alignment from language models to DiffusionDPO for images and VideoDPO for videos. PhyGDPO emphasizes physical plausibility, a real-video winner, groupwise-derived supervision, and a shared-backbone reference rather than primarily aesthetic preference alignment.

  • The controlled method comparison uses Flow-DPO, VideoDPO, and PhyGDPO on Wan2.1-1.3B with the same training data and settings.
  • The LoRA-SR ablation supports the memory argument for the evaluated configuration; it does not establish that every earlier DPO implementation requires full-model duplication.

5 Conclusion

The conclusion presents PhyAugPipe, PhyVidGen-135K, and PhyGDPO with PGR and LoRA-SR as a combined approach to physically consistent video post-training. The experiments report improved benchmark scores and human preference without LLM prompt extension at inference, with 17K resampled pairs used for the reported training run.

Appendix

  • Mathematical proof for Ineq. (8) proves the product inequality numbered Eq. (7) in the main text, despite the appendix heading’s numbering. Concavity of t raised to αj for 0 < αj ≤ 1 gives (u + v)αj ≤ uαj + vαj; setting u = 1 and v = exp(xj), then applying γj ≥ 1/αj, bounds each exp(xj) by a nonnegative aj. Expanding Πj(1 + aj) adds only nonnegative higher-order terms, establishing the group-to-pairwise upper bound.

Brief Thoughts

The strongest evidence comes from combining physics-focused data selection with real-versus-generated preference supervision and memory-efficient adaptation. Matched-setting comparisons against Flow-DPO and VideoDPO, human preferences, and a second VLM evaluator provide more persuasive support than the selected qualitative examples alone.

The main interpretive limit is the gap between improved physical plausibility and guaranteed physics learning. Real footage supplies physically grounded observations, but preferring it over generated videos can also reward appearance, capture characteristics, or prompt fit; the evaluations do not separate these signals or test physical laws through quantitative measurements. Difficult cases remain unresolved: only 0.0500 of the hard-action outputs satisfy both VideoPhy2 thresholds, and the method trails competitors on VideoPhy2 Interaction and PhyGenBench Material. Reproducibility is limited by an unspecified numerical group size m. Reporting inconsistencies include the memory-reduction percentage and prose references to Figure 9 for examples numbered Figures 7 and 8. The user-study protocol mentions 8 baselines, whereas Table 1c lists 7 competing methods. Its reference to non-repeated baseline selection is also unclear across 48 trials; clarifying the comparison allocation would make the evidence easier to audit.