Gorio Tech Blog search

Motion Attribution for Video Generation | Summary

|

Contents

This article explains the key points of Motion Attribution for Video Generation.

  • 2026-01-13 (arXiv)
  • Wu, Xindi, Paschalidou, Despoina, Gao, Jun, Torralba, Antonio, Leal-Taixé, Laura, Russakovsky, Olga, Fidler, Sanja, Lorraine, Jonathan.
  • NVIDIA, Princeton University, MIT, University of Michigan, University of Toronto, Vector Institute
  • Paper
  • Project Page

Read this article in Korean


Summary

  • Motion Attribution for Video Generation introduces Motive (MOTIon attribution for Video gEneration), a gradient-based framework for identifying fine-tuning clips relevant to specified motion patterns. It asks which candidate training clips can improve generated temporal dynamics, rather than merely sharing objects, textures, or backgrounds with a query video.
  • Motive weights latent-space prediction errors using optical-flow magnitude, computes training and query gradients with a shared timestep and noise draw, and compares normalized Fastfood projections of those gradients. Percentile-based voting aggregates rankings across queries to select fine-tuning subsets.
  • Using 10% of the candidate training data, Motive achieves VBench dynamic degree scores of 47.6% on Wan2.1-T2V-1.3B and 48.3% on Wan2.2-TI2V-5B, compared with 42.0% and 45.3% for full candidate-set fine-tuning. Human evaluators prefer its motion over the pretrained base model in 74.1% of comparisons. Motion smoothness changes little on the Wan models, and improvements are not uniform across visual-quality metrics.
  • The experiments support motion-targeted subset selection under the tested fine-tuning conditions, not exact causal attribution of individual clips. Constraints include approximately 150 hours of upfront gradient computation for 10k clips with Wan2.1-T2V-1.3B on one A100, evaluation limited to ten motion categories, imperfect separation of camera and object motion, and unmodeled classifier-free guidance.

1. Introduction

The introduction motivates attribution as a tool for understanding and curating the data that shapes trajectories, deformation, camera movement, and interactions. It focuses on fine-tuning because large pretraining corpora may be inaccessible, while selected clips can change a pretrained generator’s motion behavior.

  • Directly extending image-oriented attribution to video can emphasize static appearance; the paper identifies motion localization, sequence-level computational cost, and temporal relationships as the principal challenges.
  • The proposed contributions are scalable gradient attribution, treatment of video-length effects, motion-weighted attribution, and downstream evaluation of selected subsets.
  • Project URL: https://research.nvidia.com/labs/sil/projects/MOTIVE/

2. Background

The background connects latent-space video generation with gradient-based data attribution. Diffusion and flow matching both produce timestep-dependent gradients, while video adds variable duration and motion-specific failure modes.

2.1. Video Generation with Diffusion and Flow-Matching Models

Videos are encoded into VAE latents and corrupted according to scheduler-controlled signal and noise scales. Diffusion predicts injected noise, whereas flow matching predicts the velocity of a latent interpolant; both objectives integrate prediction errors over timesteps and noise draws.

  • Video models must address inconsistent trajectories, flicker, identity drift, and implausible dynamics even when individual frames are sharp.
  • The paper represents motion using displacement between consecutive frames; optical-flow magnitude supplies the spatial and temporal saliency used by Motive’s loss masks.

2.2. Data Attribution

Classical influence functions approximate how upweighting a training example changes a query loss using training and query gradients together with an inverse Hessian. Because curvature computation is impractical at modern model scales, methods such as TracIn and TRAK use gradient similarities or projected gradient features.

  • Diffusion gradient norms vary across timesteps, potentially biasing attribution toward large-norm evaluations; Diffusion-ReTrac motivates gradient normalization and timestep/noise subsampling.
  • Whole-video gradient similarity still mixes motion with appearance, and its cost grows with sampled timesteps, noise draws, sequence length, and gradient dimensionality.

3. Method

Motive combines scalable gradient computation, treatment of frame-length bias, motion-aware loss weighting, and influence-based subset selection. Figure 1 shows how motion extraction and gradient projection produce query-conditioned rankings without changing forward noising.

The diagram separates motion weighting from computational approximation. Optical-flow-derived masks determine which prediction errors contribute to gradients, while a shared timestep/noise draw and low-dimensional projection produce reusable training vectors for query scoring and subset selection.

Motive pipeline connecting motion extraction, loss-space weighting, projected gradients, and influence-based fine-tuning selection.
Motive pipeline connecting motion extraction, loss-space weighting, projected gradients, and influence-based fine-tuning selection.

3.1. Problem Formulation

Given a fine-tuning corpus and a query video with its conditioning, the method assigns each candidate clip a motion-aware influence score. The desired properties are predictive rankings, assessed through selected-subset fine-tuning, and feasible computation without explicit Hessian inversion or extensive per-example integration.

  • For a budget K much smaller than the corpus size N, single-query selection takes the top-K scores; multiple queries require aggregation before selecting the final subset.
  • The practical score measures gradient alignment rather than a clip’s exact counterfactual contribution.

3.2. Scalable Gradient-based Attribution

The scalable estimator uses an identity preconditioner instead of inverse-Hessian computation and evaluates training and query gradients under matched timestep/noise pairs. Its single-sample variant fixes one timestep and one shared Gaussian noise draw, reducing each example to one forward-and-backward computation.

  • Common randomness is intended to stabilize relative rankings compared with independent stochastic evaluations.
  • The single-timestep approximation is assessed against a ten-timestep reference in Section 4.4, not guaranteed to reproduce fully integrated influence.

Fastfood implements a structured Johnson–Lindenstrauss projection using Walsh–Hadamard transforms, random signs, permutation, Gaussian scaling, and rescaling. Each projected gradient is normalized, and attribution becomes its cosine similarity with the projected query gradient.

  • The reported Wan2.1 implementation projects gradients from D = 1,418,996,800 to D′ = 512 dimensions.
  • The paper reports projection complexity O(D′ log D′), dot-product complexity O(D′), and projected-vector storage O(|𝒟| D′), where |𝒟| denotes the dataset size.

3.3. Video-specific Frame-length Bias Fix

Section 3.3 attributes duration bias to raw gradient magnitudes and proposes dividing each clip’s gradient by its frame count before projection and L2 normalization. The ablation in Section 4.4 evaluates standardizing inputs to 81 frames at 16 fps, which is distinct from scalar gradient rescaling.

  • Temporal input standardization can change gradient direction, whereas positive scalar rescaling cancels under the stated linear projection followed by L2 normalization for nonzero projected gradients.
  • Figure 4 provides a qualitative duration comparison, and Section 4.4 reports the numerical length-correlation analysis.

3.4. Motion Attribution

AllTracker extracts per-pixel displacement, visibility, and confidence information. Motive constructs weights from displacement magnitude, min–max normalizes them over all frames and pixels to [0, 1] with ζ = 10⁻⁶, and bilinearly maps them to the VAE latent grid.

  • For the Wan2.1 formulation, the paper specifies spatial downsampling s = 8 and C = 16 latent channels.
  • Relative magnitude weighting emphasizes moving locations within a clip rather than ranking clips directly by their absolute amount of motion.
  • Although the tracker supplies visibility and confidence channels, the stated weighting equations use displacement magnitude.

The motion-weighted objective multiplies each latent-location squared prediction error by its motion weight, averages the weighted errors, and includes a 1/Fv factor. Motion influence is the inner product of normalized projected training and query gradients computed from this objective.

  • All-one masks recover the unmasked objective with the stated frame-count factor.
  • Masking changes the loss used for attribution, not forward noising or the generation procedure.
  • The weights suppress static-region contributions, but appearance errors within moving regions can remain entangled with temporal information.

3.5. Most Influential Fine-tuning Subset Selection

Single-query selection takes the highest-scoring clips, with the budget expressed as a dataset percentile. For multiple queries, each query votes for clips above its percentile cutoff; candidates are ranked by total votes, and the top-K are selected using the ICONS aggregation approach.

  • Voting avoids requiring raw influence scores to be calibrated across queries and favors clips repeatedly selected for different targets.
  • Figure 2 illustrates floating and rolling queries, with water dynamics and planetary rotation among positive examples and cartoons or camera-dominated footage among negative examples; these are ranking illustrations rather than individual retraining tests.

High-ranked examples connect visually different clips through query-relevant dynamics: water scenes accompany the floating query, and planetary rotation accompanies the rolling query. The negative examples include cartoons and camera-dominated footage, illustrating that gradient alignment differs from broad visual similarity without proving that each negatively scored clip would independently degrade generation.

Floating and rolling queries with positive- and negative-influence training examples from Wan2.1-T2V-1.3B attribution.
Floating and rolling queries with positive- and negative-influence training examples from Wan2.1-T2V-1.3B attribution.

3.6. Computational Efficiency Analysis

Fixing one timestep/noise pair reduces gradient computation from O(N |𝒯| B) to O(N B), where |𝒯| counts sampled pairs and B is one forward-and-backward cost. Projected-vector storage is O(N D′) rather than O(N D), while scoring one query costs O(N D′) and sorting costs O(N log N).

  • The structured Fastfood state additionally requires O(D) storage.
  • Motion masks are extracted once and cached; the paper gives extraction complexity O(N H W F).
  • Candidate-gradient computation is amortized across queries at a fixed model checkpoint, but each new query still requires its own backward pass.

4. Experiment

The experiments evaluate whether attribution-selected subsets improve generated motion using automated metrics, qualitative comparisons, human judgments, and approximation ablations. They compare selected-data fine-tuning with pretrained models, full candidate-set fine-tuning, and alternative selection methods.

4.1. Setup

The experiments use 10k-video candidate sets from VIDGEN-1M and 4DNeX-10M. Attribution targets comprise 50 query clips, with five examples for each of compress, bounce, roll, explode, float, free fall, slide, spin, stretch, and swing.

  • The appendix identifies the queries as Veo-3-generated videos manually screened for clarity, realism, and physical plausibility; they specify attribution targets and are not fine-tuning data.
  • Test prompts cover the same ten motion categories with different visual appearances, using five evaluation prompts per category.

The main models are Wan2.1-T2V-1.3B and Wan2.2-TI2V-5B, with LTX-2B evaluated in the appendix. Selection baselines include uniform random sampling, highest average motion magnitude, V-JEPA-based selection, and Motive without motion masking; subset methods use 10% of the candidate data.

  • Fine-tuning updates the DiT backbone while freezing the T5 text encoder and VAE, using 480 × 832 resolution and a learning rate of 1 × 10⁻⁵.
  • Specialists use selections for one motion category, while generalists use aggregated selections.
  • VBench measures subject consistency, background consistency, motion smoothness, dynamic degree, aesthetic quality, and imaging quality; motion smoothness and dynamic degree are the primary temporal targets.

4.2. Main Results

Table 1 shows the clearest advantage in dynamic degree: Motive obtains 47.6% and 48.3% on the two Wan models, compared with 41.3% and 41.6% for random selection, 43.8% for unmasked attribution on both models, and 42.0% and 45.3% for full fine-tuning. Its aesthetic quality scores, 46.0% and 45.6%, are also the highest among the reported methods.

  • Motion smoothness is 96.3% on Wan2.1, matching the base model and full fine-tuning, and 97.6% on Wan2.2, compared with 97.5% for both references.
  • The quantitative evidence is stronger for increasing motion activity than for improving smoothness, which is already high; dynamic degree does not directly establish physical plausibility.
  • Visual-quality trade-offs remain: Wan2.1 background consistency is 96.1% versus 96.4% for Base, and imaging quality is 64.6% versus 65.7%; Wan2.2 imaging quality is 65.5% versus 66.2% for Full FT.

At the same 10% selection budget as the other subset methods, Motive has the highest dynamic degree on both Wan models and exceeds full candidate-set fine-tuning on that metric. Motion smoothness is unchanged relative to Base on Wan2.1 and increases only slightly on Wan2.2; background and imaging scores show that the improvement is not uniform across quality dimensions.

VBench comparison across six quality dimensions for Wan2.1-T2V-1.3B and Wan2.2-TI2V-5B, with values reported as percentages.
VBench comparison across six quality dimensions for Wan2.1-T2V-1.3B and Wan2.2-TI2V-5B, with values reported as percentages.

Figure 3 compares compression, spinning, sliding, and free fall under the base model, random-selected fine-tuning, and Motive-selected fine-tuning. The displayed sequences illustrate clearer deformation and more prompt-consistent object behavior in the Motive examples.

  • Figure 2 also shows that high-ranked candidate clips need not share the query’s object identity.
  • These selected examples complement aggregate metrics but do not quantify the prevalence of improvements or directly test compliance with physical laws.

The Motive rows show a visibly flattened ball, an upright coin at the displayed spinning time points, a mug supported by the counter, and a ball at distinct descending positions. These frames illustrate more prompt-consistent behavior in the four selected examples, but cannot quantify trajectory smoothness or the frequency of improvements.

Qualitative motion comparisons between the pretrained Wan2.1-T2V-1.3B model, random-selected fine-tuning, and Motive-selected fine-tuning.
Qualitative motion comparisons between the pretrained Wan2.1-T2V-1.3B model, random-selected fine-tuning, and Motive-selected fine-tuning.

4.3. Human Evaluation

Seventeen participants judge motion in pairwise video comparisons, with randomized presentation order and ties allowed; Table 2 describes 50 videos and 850 total judgments. Motive’s win/tie/loss rates are 74.1%/12.3%/13.6% against Base, 58.9%/12.1%/29.0% against Random, 53.1%/14.8%/32.1% against Full FT, and 46.9%/20.0%/33.1% against the version without motion masking.

  • Against unmasked attribution, wins exceed losses, but the win rate is below 50% when ties remain in the denominator.
  • The protocol text refers to three pairings, whereas the table lists four baseline comparisons; the PDF does not reconcile these counts with its stated total.
  • No confidence intervals or statistical significance tests accompany the preference rates.

Preference is strongest against the pretrained base model and narrower against full fine-tuning or unmasked attribution. Against the unmasked method, 46.9% wins exceed 33.1% losses, with 20.0% ties; this is not a greater-than-50% win rate when ties are included.

Human motion-preference win, tie, and loss rates for Motive against four comparison methods.
Human motion-preference win, tie, and loss rates for Motive against four comparison methods.

4.4. Ablations

The projection ablation measures Spearman correlation between projected and full-gradient attribution rankings. Correlation rises from 46.9% at D′ = 128 to 74.7% at D′ = 512, then only to 75.7% at 1024 and 76.1% at 2048, motivating the 512-dimensional setting in Figure 5.

  • These correlations support a compact approximation but do not mean that the corresponding percentage of ranks is preserved exactly.

Spearman correlation rises rapidly up to 512 projected dimensions and then largely plateaus. The increase from 74.7% at 512 dimensions to 76.1% at 2048 supports the smaller representation as an efficiency choice, while showing that projected and full-gradient rankings are not identical.

Spearman ranking correlation with full-gradient attribution across Fastfood projection dimensions.
Spearman ranking correlation with full-gradient attribution across Fastfood projection dimensions.

At the fixed timestep t = 751, single-timestep attribution obtains a Spearman correlation of ρ = 66% with a reference using ten evenly spaced timesteps under the flow-matching schedule. The authors motivate the setting as a compromise between noisy early-denoising stages and detail-dominated late stages.

  • The reference is a multi-timestep gradient estimate, not retraining-based ground truth for individual-example influence.
  • The paper calls t = 751 the midpoint of a 1000-step trajectory without explaining that designation.

The duration ablation standardizes videos to 81 frames at 16 fps, satisfying the 4n+1 constraint. Without standardization, attribution scores reportedly correlate with video length at ρ = 78.0%; standardization reduces spurious length correlations by 54.0%, while Figure 4 illustrates more motion-relevant floating-query selections.

  • Among Figure 4’s ranking panels, “Without Frame Normalization” is on the left and “With Frame Normalization” is on the right, opposite the caption and Section 4.4’s prose.
  • The experiment supports temporal input standardization but does not isolate an effect of scalar gradient division before cosine normalization.

The right-hand ranking panel labeled “With Frame Normalization” contains water-related footage, whereas the panel labeled “Without Frame Normalization” contains less motion-coherent selections for the floating query. The caption and Section 4.4 reverse these panel positions. The example illustrates the reported benefit of temporal standardization but does not isolate scalar gradient rescaling.

Floating-query rankings with and without frame normalization, identified by the panel labels.
Floating-query rankings with and without frame normalization, identified by the panel labels.

5. Conclusion

The conclusion presents motion-aware attribution as a data-selection approach to improving video dynamics through targeted fine-tuning. The experiments support smaller, motion-relevant subsets, while broader diagnostic and controllability applications are proposed uses rather than separately validated outcomes.

Appendix

  • A. Notation: Tables 3 and 4 define the video, conditioning, latent-space, dataset, motion, loss, projection, and influence symbols. They distinguish full gradients from projected gradients and specify the fixed timestep/noise pair used for attribution.
  • B. Extended Related Work, B.1. Data Attribution, and B.2. Motion in Video Generation: The appendix distinguishes retraining-based attribution from gradient-based approximations, reviews timestep-dependent diffusion attribution, and situates Motive alongside temporal modeling and optical-flow methods. Motive applies motion-weighted gradients to candidate training-clip selection rather than changing the generator architecture or providing motion-control signals.
  • C. Additional Experiments and C.1. Results on Additional Video Generation Models: Table 5 extends evaluation to LTX-2B. Motive reaches 45.1% dynamic degree versus 36.5% for Base, 38.9% for Full FT, and 41.8% without motion masking; motion smoothness reaches 95.5% versus 94.6%, 94.8%, and 95.3%, respectively. Its imaging quality of 48.6% remains below Full FT at 49.1%, and its aesthetic quality of 32.7% is below the 32.8% reported for random selection and unmasked attribution.
  • D. Analysis, D.1. Motion Distribution Analysis, and D.2. Cross-Motion Influence Patterns: Figure 6 analyzes 10k VIDGEN clips and reports mean motion magnitudes of 3.85 for the top 10% and 3.69 for the bottom 10%, a 4.3% relative difference; both groups span low- and high-motion bins. Figure 7 reports mean cross-category influential-set overlaps of 24.0% on 4DNEX and 24.3% on VIDGEN, with bounce-to-float overlap of 44.4%/46.3% and free fall-to-stretch overlap of 12.8%/12.7%. The paper attributes directional asymmetry to differences in the number of unique influential clips aggregated across each category’s five queries. These analyses distinguish the selections from a simple highest-motion filter, but selected-clip overlap does not establish a shared internal motion mechanism.
  • E. Additional Method Details: Algorithm 1 summarizes motion-mask extraction, motion-weighted gradient computation, projection, cosine scoring, and subset selection. The authors describe the tracker as a replaceable saliency source and the attribution procedure as applicable to both diffusion and flow-matching objectives. Algorithm 1 divides gradients by frame length after evaluating a loss that already contains a frame-count factor; the PDF does not clarify why both scalings are listed.
  • F. Additional Experiment Details, F.1. Hyperparameter Settings, Samples from Query Set, Samples from Test Set, and F.2. Details on Motion Query Data: The appendix specifies bfloat16 attribution, 512-dimensional projections, AdamW fine-tuning, top-10% selection, and training for 1 epoch while repeating the dataset 50 times. Defaults come from DiffSynth-Studio and the official Wan repository for Wan models, and Diffusion-Pipe for LTX models. Query and test examples change appearance within a motion category, including sandwich bread versus a rubber ball for compression and a foam cube versus a leaf for floating. Figure 8 illustrates manually screened Veo-3 queries designed to reduce background and camera-motion confounds; they specify attribution targets and are not used for fine-tuning.
  • G. Discussion, G.1. Runtime, G.2. Limitations, and G.3. Future Directions: Tables 6 and 7 report approximately 150 hours of candidate-gradient computation for 10k clips with Wan2.1-T2V-1.3B on one A100, versus approximately 3 hours for V-JEPA selection and 5.5 hours for motion magnitude. Candidate gradients are reusable across queries at a fixed checkpoint; a new query requires approximately 54 seconds for its gradient and 46 milliseconds for similarity computation. Projection costs approximately 1.97 seconds per sample, and aggregation across 50 queries costs approximately 139 milliseconds. Limitations include dilution of informative motion segments, camera-motion sensitivity, unmodeled classifier-free guidance, and trade-offs in broader capabilities. The authors state that spatially uniform motion masks help identify and down-weight camera-only clips, but do not specify the rule. Proposed extensions include segment-level attribution, confidence-aware masks, self-generated queries, iterative curation, learned query weights, other modalities, and safety-oriented filtering and auditing.
  • H. Visualization and H.1. Motion Visualization: Figure 9 presents seven clips at early, middle, and late time points, pairing original frames with motion overlays that attenuate static regions toward gray. The overlays visualize loss weights rather than internal model representations, and appearance information remains visible in moving regions.

Brief Thoughts

The comparisons with unmasked attribution and motion-magnitude selection provide the strongest evidence for motion-conditioned data curation: Motive improves dynamic degree across all three tested models, while the distribution analysis shows that selected clips are not simply those with the most movement. The results support targeted selection within the tested motion categories, but do not establish general physical realism or uniformly better generation quality.

Gradient cosine similarity is an influence proxy validated through subset fine-tuning, not an exact causal explanation of individual clips. Under the stated linear projection and L2 normalization, positive frame-count rescaling cancels for nonzero projected gradients, so the duration results should be attributed to temporal input standardization unless further implementation evidence establishes a separate effect.