4DAnyone: Create Anyone in 4D from a Casual Monocular Video | Summary
20 Aug 2026 | Paper Review 4D Human Reconstruction Multiview Video Generation Video Diffusion Models 4D Gaussian SplattingContents
This article explains the key points of 4DAnyone: Create Anyone in 4D from a Casual Monocular Video.
- 2026-08-20 (arXiv)
- Jin, Yudong, Xie, Tao, Zhang, Qihang, Shen, Zehong, Xu, Zhen, Shen, Yujun, Bao, Hujun, Zhou, Xiaowei, Xu, Yinghao.
- State Key Lab of CAD&CG, Zhejiang University, China, Robbyant, China, Chinese University of Hong Kong, China, Ant Group, China, Hong Kong University of Science and Technology, China
- Paper
- Github
- Project Page
Summary
- 4DAnyone reconstructs a free-viewpoint 4D human from an uncalibrated casual monocular video by generating reconstruction-grade multiview videos and fitting FreeTimeGS. Its central scaling problem is bounded DiT attention: generating the tens of views needed for 4D Gaussian Splatting (4DGS) requires grouping target views, while accumulated reference views otherwise make conditioning grow as (O(N)) and disjoint target groups can drift structurally.
The overview depicts the intended pipeline: a casually captured monocular clip is expanded into multiview-consistent target videos, which are used to construct a free-viewpoint volumetric video.
1 Introduction
Direct monocular avatar reconstruction struggles with occluded appearance and generally relies on camera parameters, while camera-controlled diffusion models can synthesize plausible novel views without maintaining consistency at reconstruction-scale view counts. The paper uses sparse 3D skeleton conditioning rather than dense depth or source-camera parameters, and introduces Reference Context Packing (RCP) and Target Context Routing (TCR) to address appearance conditioning and cross-group generation consistency.
3 Method
Given a source video Vsrc with unknown intrinsics and poses, the method generates videos at prescribed static viewpoints and reconstructs them with FreeTimeGS. GVHMR estimates a 3D skeleton sequence; depth-buffered target-view skeleton renderings condition a Wan2.2-based DiT, and its generated videos serve as multiview observations for 4DGS.
3.1 Overview
The architecture encodes the source video, packed reference context, noisy target latents, and skeleton conditions. The DiT combines video attention, multiview attention, and text cross-attention; generated videos are progressively incorporated into the reference context, while TCR coordinates denoising across target-view groups within a fixed per-group memory budget.
The diagram traces the path from HMR-derived depth-buffered skeletons through the skeleton encoder and diffusion transformer to target videos, and shows that those videos subsequently train a FreeTimeGS reconstruction.
3.2 3D-Aware Skeleton Conditioning
Recovered 3D skeletons are rasterized with a pixelwise z-buffer so nearer limbs occlude farther ones, resolving the front-back ambiguity of ordinary 2D pose renderings without extra input channels. The method retains 40 Goliath keypoints—17 body, 6 foot, 10 palm-level hand, and 7 auxiliary landmarks—and excludes facial points and individual finger joints; a zero-initialized skeleton encoder adds residual tokens to noisy latents.
3.3 Scalable Multiview Consistency
RCP compresses accumulated generated-reference videos into a fixed token budget with 2× and 4× multi-scale patchify layers, changing reference-context growth from (O(N)) to (O(1)). TCR cyclically regroups four target views during high-noise steps to communicate global structure, then fixes adjacent four-view groups at low noise to refine local detail; the default switching ratio is (t_s/T = 0.2).
The figure separates the two scaling mechanisms: RCP maintains a fixed-size context for accumulated reference views, while TCR changes target-group membership early in denoising and keeps adjacent groups fixed later.
3.4 Training Protocol
Training follows a three-stage curriculum: foreground-only DNA-Rendering teaches skeleton-conditioned camera control, unmasked multiview data teaches backgrounds and lighting, and Pexels plus TedTalk monocular videos improve in-the-wild generalization. The model is fine-tuned from Wan2.2-TI2V-5B at 704×1280 with latent flow-matching MSE and LPIPS weighted by (\lambda = 0.25); Stage 3 removes finger keypoints.
4 Experiments
Experiments evaluate direct generated-video reconstruction, consistency among generated views, and downstream 4DGS reconstruction on held-out DNA-Rendering scenes and out-of-distribution DyMVHumans scenes. The study compares MV-Performer, TrajectoryCrafter, and a ReCamMaster variant fine-tuned with the same datasets, RCP, and TCR; all three evaluation dimensions use PSNR, SSIM, and LPIPS.
4.1 Experimental Setup
Training data comprise MVGameHuman, SynCamVideo, DNA-Rendering, TedTalk, and Pexels. Quantitative testing uses 10 DNA-Rendering and 3 DyMVHumans scenes, 16 approximately uniformly distributed target cameras, and 98 frames per scene from a single front-view source; inference uses 20 denoising steps and four-view target groups.
The table establishes the scale and composition of the training corpus. Three datasets provide multiview coverage, while TedTalk and Pexels contribute monocular video diversity for generalization.
4.2 Comparison with State-of-the-Art
4DAnyone is best in every reported Table 2 metric. For 4DGS reconstruction, it reaches 24.15 PSNR, 0.863 SSIM, and 0.159 LPIPS on DNA-Rendering, and 23.28, 0.846, and 0.117 on DyMVHumans; generated-video consistency is 24.33/0.862/0.163 and 24.48/0.862/0.109, respectively. The qualitative comparisons associate baseline failures with side- and back-view distortions, depth-error accumulation, or inaccurate camera control.
The table supplies the quantitative comparison across generated-video consistency, downstream 4DGS reconstruction, and direct generated-video reconstruction. 4DAnyone has the best reported values in every DNA-Rendering and DyMVHumans column.
4.3 Ablation Study
Removing both RCP and TCR yields 21.09 PSNR, 0.766 SSIM, and 0.216 LPIPS in the DNA-Rendering generated-video-consistency ablation, compared with 22.63, 0.796, and 0.191 for the full Sliding configuration. Each component improves the result, Sliding routing outperforms the reported Random and Strided alternatives, and depth-buffered skeletons visually resolve ambiguities present with 2D skeleton conditioning. The supplementary switching-time sweep identifies (t_s/T = 0.2) as the start of a saturation regime, although several lower ratios produce closely comparable scores.
The left example illustrates that depth-aware skeleton rendering resolves a side-view body-part ambiguity seen with 2D skeleton conditioning. The right example compares two independently grouped target views and shows more consistent appearance when RCP and TCR are enabled.
4.4 Generalization to In-the-Wild Videos
The paper presents qualitative results on diverse in-the-wild videos, including back-view sources, complex appearances, and complex motion, showing generated target views and 4DGS renderings. These examples support qualitative generalization beyond the controlled benchmarks, but they are not a separate quantitative wild-video benchmark.
The examples test difficult source conditions outside the standard front-view setting, including a back-view source, detailed appearance, and complex motion. They provide qualitative evidence for robustness of target-view generation.
Each row places a source video beside four of sixteen generated views and a 4DGS rendering for a real-world human video. The figure is the paper's primary visual evidence for in-the-wild generalization.
Appendix
- The supplementary material specifies a 10-layer Conv3d skeleton encoder, MVGameHuman with 38k synchronized 24-camera videos of 318 actors, and a 3.5 mm held-out keypoint error for the SMPL-X-to-Goliath70 regressor. A 16-camera, 121-frame reconstruction is reported to require approximately 2 min of preprocessing on one RTX 4090, ∼7 min to generate four videos on one H20 GPU, and ∼30 min of FreeTimeGS optimization on one RTX 4090. Limitations include inconsistent views for loose garments moving far from the body and propagation of HMR pose errors into generated views.
Brief Thoughts
The paper identifies cross-group drift as a practical obstacle to applying multiview diffusion at the view counts needed for 4D reconstruction, and addresses it through context packing and inference-time routing without requiring calibrated source cameras or dense depth. Evidence is strongest on the controlled DNA-Rendering and DyMVHumans evaluations and their component ablations; performance remains dependent on reliable human motion recovery and on clothing motion being represented adequately by the body skeleton. The paper also notes deepfake misuse, identity privacy, and copyright risks of realistic human-video synthesis.