4DAnyone: Create Anyone in 4D from a Casual Monocular Video | Summary
20 Aug 2026 | Paper Review 4D Human Reconstruction Multiview Video Generation Video Diffusion Models 4D Gaussian SplattingContents
- Summary
- 1 Introduction
- 2 Related Work
- 3 Method
- 4 Experiments
- 5 Conclusion
- Acknowledgments
- A Model Details
- B Dataset Details
- C Training Details
- D HMR Details
- E Inference Details
- G Additional Results
- H Limitations and Failure Cases
- Appendix
- Brief Thoughts
This article explains the key points of 4DAnyone: Create Anyone in 4D from a Casual Monocular Video.
- 2026-08-20 (arXiv)
- Jin, Yudong, Xie, Tao, Zhang, Qihang, Shen, Zehong, Xu, Zhen, Shen, Yujun, Bao, Hujun, Zhou, Xiaowei, Xu, Yinghao.
- State Key Lab of CAD&CG, Zhejiang University, China, Robbyant, China, Chinese University of Hong Kong, China, Ant Group, China, Hong Kong University of Science and Technology, China
- Paper
- Github
- Project Page
Summary
- 4DAnyone reconstructs a free-viewpoint 4D human representation from an uncalibrated monocular video. It generates reconstruction-grade multiview videos conditioned on recovered 3D skeletons, then uses FreeTimeGS to lift them into 4D Gaussian Splatting (4DGS); Reference Context Packing (RCP) and Target Context Routing (TCR) address consistency when generation must scale to many target views.
The diagram establishes the intended pipeline: a casually captured monocular human video is expanded into multiview videos and then reconstructed as a free-viewpoint volumetric sequence. It motivates the requirement for consistency across generated views rather than plausible views in isolation.
1 Introduction
The paper targets the gap between casually captured monocular human video and the dense calibrated multiview observations typically required for 4DGS. When target views exceed one diffusion transformer’s practical attention budget, splitting them into groups causes growing reference-context cost and prevents direct cross-group communication, producing appearance inconsistency and structural drift.
2 Related Work
Prior camera-controlled video generation is divided into implicit camera conditioning and explicit dense-geometry conditioning. The former can drift under large viewpoint changes, while the latter depends on depth or 3D coordinates that are difficult to estimate reliably from unconstrained video; pose-driven animation generally emphasizes motion transfer rather than reconstruction-grade novel-view synthesis.
3 Method
Given a source video Vsrc ∈ R^{F×3×H×W} with unknown intrinsics and poses, the method estimates a 3D skeleton sequence and renders it from prescribed target viewpoints. A Wan2.2-based DiT processes source tokens, packed reference tokens, and skeleton-conditioned target tokens with video attention, multiview attention, and text cross-attention; FreeTimeGS then reconstructs the final 4DGS model.
3.1 Overview
The pipeline generates target views around the subject under a fixed per-group memory budget. HMR provides target-view skeleton conditions, RCP supplies bounded appearance context from generated references, and TCR changes target-group membership during denoising to communicate structure across groups.
The architecture traces data flow from HMR-derived target-view skeletons through the skeleton encoder and diffusion transformer to generated videos and 4DGS training. It shows that generated views are repacked as reference context for later generation rounds.
3.2 3D-Aware Skeleton Conditioning
4DAnyone uses sparse 3D skeletons rather than estimated dense depth, following an accuracy-over-density principle. It rasterizes skeletons with a pixelwise z-buffer so nearer limbs occlude farther limbs, and selects 40 of the 308 Goliath keypoints: body, foot, palm-level hand, neck, shoulder, and elbow landmarks, excluding facial points and individual finger joints.
3.3 Scalable Multiview Consistency
RCP packs generated references at mixed spatial resolutions into fixed reference-token slots, changing reference-context growth from O(N) to O(1). It uses compression ratios r ∈ {2, 4} and fixed context CR = [P1(Vsrc), {P2(Vaj)}^3_{j=1}, {P4(Vbj)}^4_{j=1}]; progressive inference first generates reference views using farthest-point sampling, then reuses this context for target groups.
The figure separates the two scalability mechanisms during progressive inference. RCP packs references into a fixed context, while TCR cyclically regroups four-view target sets at high noise and fixes adjacent groups at low noise.
3.4 Training Protocol
Training uses a three-stage curriculum: foreground-only DNA-Rendering for skeleton-conditioned camera control, unmasked multiview data for background, lighting, and shadow modeling, and monocular Pexels and TedTalk data for in-the-wild generalization. The objective is L = Llatent + λLLPIPS with λ = 0.25, with LPIPS evaluated on body-part-aware crops to reduce memory use.
4 Experiments
Experiments assess generated-video consistency, 4DGS reconstruction from generated views, and direct generated-video reconstruction against ground truth. Comparisons include MV-Performer, TrajectoryCrafter, and ReCamMaster†, where † denotes an author-fine-tuned ReCamMaster equipped with the same RCP and TCR strategy.
4.1 Experimental Setup
The model is built on Wan2.2-TI2V-5B and trained at 704×1280 with learning rate 1 × 10−5 on 128 H20-3E GPUs; the three stages take approximately 0.5, 1, and 1.5 days. Quantitative evaluation uses 10 unseen DNA-Rendering scenes and 3 out-of-distribution DyMVHumans scenes, with 16 approximately uniform cameras and 98 frames per scene; inference uses 20 denoising steps, four-view groups, and ts/T = 0.2.
The table records the scale and modality of the training corpus. MVGameHuman provides 38k videos from 24 cameras and 318 actors, while Pexels provides 20k monocular videos spanning 1,411 actors.
4.2 Comparison with State-of-the-Art
4DAnyone is best in every reported PSNR, SSIM, and LPIPS column. On DNA-Rendering, generated-video consistency is 24.33 PSNR, 0.862 SSIM, and 0.163 LPIPS, while 4DGS reconstruction is 24.15, 0.863, and 0.159; on DyMVHumans, the respective scores are 24.48/0.862/0.109 and 23.28/0.846/0.117.
The qualitative comparison pairs target-view generation with corresponding 4DGS renderings. It illustrates the reported geometric distortion in MV-Performer and camera-control misalignment in fine-tuned ReCamMaster† relative to the proposed method.
4.3 Ablation Study
Removing RCP or TCR degrades generated-video consistency on the DNA-Rendering ablation. Full Sliding routing reaches 22.63 PSNR, 0.796 SSIM, and 0.191 LPIPS, compared with 22.03/0.780/0.203 without RCP and 22.21/0.788/0.196 without TCR; sliding routing outperforms Random and Strided routing in the reported setting.
The ablation separates RCP and TCR and compares routing strategies. Combining both components with Sliding routing gives the strongest reported generated-video-consistency result, while Random and Strided routing do not improve on it.
The left panel compares plain and depth-buffered skeleton conditioning for overlapping limbs. The right panel depicts reduced appearance disagreement between target groups when RCP and TCR are enabled.
4.4 Generalization to In-the-Wild Videos
Qualitative in-the-wild examples include back-view inputs, complex appearances, and complex motions, followed by 4DGS renderings. They demonstrate the intended behavior beyond the controlled benchmarks, but do not constitute an additional quantitative in-the-wild evaluation.
Each row links a source video to four of sixteen generated target views and a 4DGS novel-view rendering. The gallery makes the generation-to-reconstruction output inspectable across varied in-the-wild videos.
5 Conclusion
The conclusion attributes reconstruction-grade multiview generation to sparse 3D skeleton conditioning, RCP, and TCR. The paper also identifies deepfake misuse, privacy violations, and copyright infringement as risks, and advocates disclosure of synthetic outputs and consent from depicted subjects.
Acknowledgments
The work was partially supported by the National Key R&D Program of China (No. 2024YFB2809105), NSFC (No. U24B20154), Zhejiang Provincial Natural Science Foundation of China (No. LR25F020003), Ant Group, and the Information Technology Center and State Key Lab of CAD&CG, Zhejiang University.
A Model Details
The supplementary model details rearrange multiview-attention tokens to (f, v·h·w, d), allowing viewpoints at the same timestep to attend directly. The 3D-aware skeleton encoder has 10 Conv3d layers followed by a zero-initialized 1×1×1 projection, while 2× and 4× RCP patchify layers produce one quarter and one sixteenth as many tokens as the standard patchify layer.
B Dataset Details
MVGameHuman contains 38k synchronized multiview videos rendered at 2560×1440, covering 318 actors and 24 virtual cameras per sequence. Together with SynCamVideo, DNA-Rendering, TedTalk, and Pexels, it supplies multiview, real-capture, and monocular data for the training curriculum.
C Training Details
All stages fine-tune Wan2.2-TI2V-5B at 704×1280 with λ = 0.25. Training configurations keep target-camera and frame-count products approximately constant, such as 6 × 41 ≈ 4 × 61 ≈ 1 × 121; Stage 3 removes finger keypoints because monocular finger detections are noisy.
D HMR Details
For monocular input, GVHMR estimates a ground-aligned SMPL-X mesh sequence, and a sparse vertex-to-keypoint regressor extracts Goliath keypoints for z-buffered rendering. On held-out scenes, the regressor reports 3.5 mm mean keypoint error versus 14.0 mm for a nearest-vertex baseline.
E Inference Details
The paper reports approximately 2 min for GVHMR inference and skeleton rendering on one RTX 4090, approximately 7 min to generate four 121-frame videos with 20 denoising steps on one H20 GPU, and approximately 30 min for FreeTimeGS optimization on one RTX 4090. For 16 cameras and 121 frames, FreeTimeGS is optimized for 50k iterations using Adam at 1.6 × 10−4.
G Additional Results
Additional results chain Wan-Animate with 4DAnyone: an input image and driving motion video first yield a monocular source video, which 4DAnyone converts into multiview videos for FreeTimeGS reconstruction. This is an assembled application of an off-the-shelf animation model and the proposed pipeline rather than a separately benchmarked method.
H Limitations and Failure Cases
The stated failure modes are garments moving far from the body and inaccurate HMR pose estimation. Flowing fabric is weakly constrained by a body skeleton and can become inconsistent across views, while an erroneous pose estimate, such as flat feet for an en-pointe dancer, is propagated to generated views.
Appendix
- The supplementary material details the architecture, datasets, training, HMR, inference, evaluation protocol, routing ablations, additional results, and failure cases. It documents implementation choices underlying RCP and TCR and clarifies the pipeline’s dependence on reliable skeleton guidance.
Brief Thoughts
The paper identifies two distinct scaling problems in reconstruction-scale multiview diffusion and assigns them to separate mechanisms: RCP for bounded reference context and TCR for cross-group target communication. Its evidence spans direct video fidelity, consistency, and downstream 4DGS evaluation, although the ReCamMaster† baseline is fine-tuned while MV-Performer and TrajectoryCrafter are evaluated zero-shot, and failure cases remain tied to clothing dynamics and HMR accuracy.