Lyra 2.0: Explorable Generative 3D Worlds | Summary
14 Apr 2026 | Paper Review 3D Scene Generation Long-Horizon Video Generation Spatial Memory Self-AugmentationContents
- Summary
- 1. Introduction
- 2. Related Work
- 3. Preliminaries
- 4. Method
- 5. Experiments
- 6. Discussion
- Appendix
- Brief Thoughts
This article explains the key points of Lyra 2.0: Explorable Generative 3D Worlds.
- 2026-04-14 (arXiv)
- Shen, Tianchang, Bahmani, Sherwin, He, Kai, Srinivasan, Sangeetha Grama, Cao, Tianshi, Ren, Jiawei, Li, Ruilong, Wang, Zian, Sharp, Nicholas, Gojcic, Zan, et al.
- NVIDIA
- Paper
Summary
- Lyra 2.0: Explorable Generative 3D Worlds generates camera-controlled videos from a single image and reconstructs them into explicit 3D environments. It addresses spatial forgetting when previously observed regions leave the context window and temporal drifting as autoregressive synthesis errors accumulate.
- The method stores independent per-frame geometry to retrieve relevant history images and guide attention through dense geometric correspondences, rather than rendering a fused global reconstruction as spatial-memory conditioning. FramePack and lightweight self-augmentation training reduce drift, while a reconstruction model fine-tuned on generated videos produces compact 3D Gaussians and surface meshes.
- On DL3DV and Tanks and Temples, the full model improves appearance metrics over the evaluated baselines and achieves the best listed reconstructed-scene metrics, although GEN3C retains better camera control and reprojection error. A 4-step distilled model reduces generation time from approximately 194 seconds to 15 seconds per 80-frame segment on a single NVIDIA GB200 GPU, with dataset-dependent quality and consistency trade-offs. The framework targets static environments and can inherit exposure inconsistencies from training data.
1. Introduction
Generative reconstruction replaces dense real-world capture with generated novel views, then lifts those views into explicit geometry. Expanding this approach across rooms or along streets requires preserving scene identity under substantial viewpoint changes and revisits, beyond consistency between neighboring frames.
- Spatial forgetting forces a finite-context model to invent previously observed structures again; temporal drifting progressively changes color, texture, and geometry through autoregressive error accumulation.
- A fused 3D memory can amplify reconstruction errors through rendered conditioning, whereas history images with camera embeddings alone leave multi-view correspondence inference to self-attention.
- Figure 1 illustrates indoor lookbacks, an outdoor example spanning approximately 90 meters, diverse visual domains, and simulation-oriented assets. The project URL is https://research.nvidia.com/labs/sil/lyra2/.
The indoor sequence includes a lookback to the starting room, while the outdoor example expands along a trajectory marked approximately 90 meters. Paired point-cloud views and 3DGS renderings connect video exploration to explicit scene assets. The diverse-domain and simulation examples demonstrate breadth qualitatively rather than measuring generalization or physical validity.
2. Related Work
The related work covers camera-conditioned video generation, memory-aware long video generation, and 3D scene generation. Lyra 2.0 combines image retrieval with geometric correspondence conditioning while avoiding a single fused geometry representation for spatial memory.
- Camera-control methods use pose vectors, Plücker rays, discrete actions, or structured 3D renderings; correspondence-based conditioning in GenWarp is restricted to single-image diffusion.
- Long-video memory methods retrieve past observations, accumulate explicit 3D scene representations, or retain internal latent states and key–value caches. FramePack compresses history into contextual slots according to temporal proximity.
- Previous generative reconstruction systems combine synthesized views with feed-forward reconstruction but typically offer limited view coverage. Lyra 2.0 progressively expands scenes through long camera trajectories.
3. Preliminaries
The video generator uses a DiT in the latent space of the Wan 2.1 causal VAE, with 8× spatial downsampling, 4× temporal downsampling, and 16 latent channels. Flow matching forms (z_t=(1-t)z_0+t\epsilon) and trains the network to predict the velocity (\epsilon-z_0).
- For F input frames, the latent temporal length is F′ = ⌊(F − 1)/4⌋ + 1; the first frame is encoded independently, while subsequent frames are temporally compressed.
- Depth-warped conditioning from the most recent frame provides camera control, but loses its signal when a large viewpoint change leaves no projected pixels in the target view. Per-pixel 6D Plücker rays provide complementary camera information.
- The FramePack layout is f1k1 as the initial-image anchor, followed by f16k4 f2k2 f1k1 as temporal slots and g20 as the generation target. Here fnkm denotes n latent frames with spatial subsampling factor m; g20 corresponds to the 80-frame video segment used in the implementation.
4. Method
The method combines long-horizon video generation with explicit 3D reconstruction. Within spatial memory, geometry identifies and aligns useful observations, while the video model synthesizes appearance from the retrieved images and its learned prior.
4.1. Overview
Starting from an input image and a user-defined camera trajectory, Lyra 2.0 runs an autoregressive retrieve–generate–update loop. Figure 2 shows how retrieved spatial context and compressed temporal history condition each new segment, after which generated frames and estimated geometry extend the memory.
- Users can supply an optional text prompt to guide outpainting and plan subsequent trajectories through an interactive 3D explorer.
- The growing memory supports revisits outside the temporal context window, while the generated video is reconstructed into 3D Gaussians or meshes.
- The architecture permits memory to grow across successive segments; the finite-horizon experiments do not establish the paper’s claimed consistency over arbitrarily long trajectories.
The diagram distinguishes the growing spatial memory from the video generator and final reconstruction. Retrieved images carry appearance, warped canonical coordinates guide correspondence-based attention, and temporal slots supply recent history. This separation allows revisits to use observations outside the temporal window without rendering fused global geometry for spatial-memory conditioning.
4.2. Anti-Forgetting for 3D-Persistent Video Generation
Building the 3D Cache stores each frame’s full-resolution depth and camera parameters alongside a downsampled world-space point cloud. Full-resolution depth supports correspondence computation, while the smaller point cloud supports retrieval.
- Per-frame point clouds are never fused into a single global point cloud for spatial memory, avoiding accumulated cross-view misalignment in a shared reconstruction.
- The design keeps imperfect depth estimates from generated images separate instead of consolidating their errors into the geometry used for future conditioning.
Geometry-Aware Retrieval ranks history observations by their visible 3D content in the target view. Points are projected into the target image plane and checked against the minimum projected depth across frames; a point is visible only when its depth difference is below the occlusion threshold.
- Training samples history frames in proportion to their visible-point counts, exposing the generator to different retrieval outcomes.
- Inference greedily selects frames that cover the most still-uncovered target pixels, reducing redundant viewpoints and maximizing collective coverage.
- The system retrieves N_s = 5 frames, uses point-cloud subsampling factor d = 8, and sets δ = 0.1 in normalized depth units.
Injecting Spatial Memory into the Video Model independently VAE-encodes retrieved images and places them in spatial slots alongside the temporal slots. Canonical coordinate maps encode normalized pixel position and retrieved-frame identity, then are forward-warped with depth to establish target-to-history correspondences without carrying RGB appearance.
- The complete layout is f1k1 as the anchor, f4k2 f1k1 as spatial slots, f16k4 f2k2 f1k1 as temporal slots, and g20 as the generation target. Missing retrieved frames are padded.
- Each canonical coordinate map has channels (u, v, 2j/N_s − 1); warped depth adds a fourth channel. Learned embeddings guide self-attention through query and key injection in every DiT block, leaving value projections unchanged.
- Coordinate warping avoids directly feeding RGB warping artifacts into spatial-memory conditioning. The separate camera-control module still depth-warps the most recent RGB frame.
- Figure 9 supports the five-slot choice: mean target-view coverage is 42.6% with five retrieved frames and 45.0% with ten, indicating diminishing coverage gains.
Mean coverage rises from 28.9% with one retrieved frame to 42.6% with five and 45.0% with ten, with a broad 25th–75th percentile band. The ground-truth-depth test supports diminishing coverage gains beyond five slots, but does not directly measure generated-video quality or retrieval runtime.
4.3. Anti-Drifting for Long-Horizon Video Generation
Self-Augmentation Training addresses observation bias: training normally conditions on clean history, whereas inference conditions on imperfect generated history. With probability p_aug = 0.7, the method samples t uniformly from [0, 0.5], corrupts the history latent, and performs one-step denoising to obtain imperfect conditioning from the model’s own prediction.
- The corrupted history is z_t^hist = (1 − t)z_0^hist + tε, and its approximate reconstruction is z̃_0^hist = z_t^hist − t·v_θ(z_t^hist, t, c).
- The current-chunk target is always encoded using the clean history’s causal VAE cache. The model must therefore recover a clean target despite receiving degraded DiT conditioning.
- This requires one additional DiT forward pass rather than full multi-step generation of each bidirectional history segment. FramePack extends temporal context, while self-augmentation trains recovery from imperfect histories.
4.4. 3D Reconstruction
3D Reconstruction adapts Depth Anything v3 (DAv3) to the resolution and residual inconsistencies of generated videos. Its Gaussian DPT head is downsampled by k × k, reducing the predicted Gaussian count by k² while retaining high-resolution image inputs; fine-tuning on generated scenes improves tolerance to generative artifacts.
- The implementation uses k = 2, giving a 4× reduction in Gaussian count.
- Surface Mesh Extraction uses an OpenVDB-based hierarchical sparse grid, with finer cells near generation viewpoints and coarser cells in distant regions.
- Median Gaussian depth and depth-gradient normals yield oriented points for a signed distance function. Marching cubes extracts surfaces, which are stitched across hierarchy levels and decimated.
4.5. Distilled Model for Accelerated Inference
Distribution Matching Distillation trains a student to generate each segment in 4 denoising steps rather than 35. Classifier-free guidance is also distilled into the student, eliminating separate conditional and unconditional passes, while self-augmentation is retained during distillation.
- The combined changes reduce segment-generation time by roughly 13×.
- The distilled model accelerates segment-based exploration; the reported latency and quality trade-offs are detailed in the experiments and appendix.
5. Experiments
The experiments evaluate generated long videos and the explicit 3D scenes reconstructed from them. Ablations test the memory and anti-drifting components, while application examples demonstrate interactive exploration and simulator export.
5.1. Training Details
Training uses DL3DV, containing 10K long real-world video clips, with camera poses estimated by ViPE, depth predicted by Depth Anything V3, and captions generated by Qwen3-VL-8B-Instruct. Each video contributes 1,000 sampled frames for conditioning–target construction.
- With 30% probability, image-to-video training generates the first L = 80 consecutive frames from one initial image.
- With 70% probability, autoregressive training samples s ∈ [0, S_max), where S_max = ⌊(T − 1)/L⌋ − 1. History covers [0, s·L + 1), and the target contains the next L consecutive frames in segment s + 1.
5.2. Evaluation on Long Video Generation
Long-video evaluation uses DL3DV-Evaluation for in-domain testing and Tanks and Temples for out-of-domain testing. Baselines are Yume-1.5, GEN3C, Context as Memory (CaM), VMem, SPMem, and HY-WorldPlay; CaM and SPMem are re-implemented using Wan2.1-14B because their implementations were not open-sourced.
- SSIM, LPIPS, and FID evaluate image fidelity and appearance distribution. WorldScore supplies Subjective Quality, Style Consistency between first and last frames, and Camera Controllability.
- Reprojection error uses per-frame depth estimated by an off-the-shelf SLAM system to assess geometric consistency.
- Yume-1.5 and HY-WorldPlay lack explicit camera-trajectory conditioning. Their Camera Controllability entries are not reported in Table 1, limiting direct comparison of control accuracy.
Table 1 shows that the full model improves the principal appearance metrics over the evaluated baselines on both datasets. On DL3DV it obtains SSIM 0.388, LPIPS 0.498, FID 43.43, Subjective Quality 44.54, and Style Consistency 87.46; on Tanks and Temples the corresponding values are 0.384, 0.552, 51.33, 43.35, and 85.07.
- GEN3C remains stronger on camera control and reprojection error: its DL3DV scores are 69.54 and 0.068, compared with 64.67 and 0.076 for the full model; on Tanks and Temples they are 70.91 and 0.054, compared with 63.87 and 0.069.
- CaM and SPMem provide competitive appearance quality but weaker camera control. The results support improved appearance with useful control, not superiority on every metric.
- Figure 3 compares frames at approximately frame 800+, where the selected baseline examples show structural distortions, collapse, or appearance drift and Lyra 2.0 retains recognizable scene structure.
Across DL3DV and Tanks and Temples, the full model leads the evaluated baselines on appearance and style metrics, but GEN3C has better camera control and lower reprojection error. The DMD rows reveal dataset-dependent changes, including a Tanks and Temples Style Consistency decrease from 85.07 to 78.91 despite slightly better FID and LPIPS.
Selected lighthouse and indoor scenes at approximately frame 800+ expose structural distortion, appearance drift, and severe degradation in baseline outputs. Lyra 2.0 retains recognizable architecture and scene appearance in these examples, although its generated views do not exactly match ground truth.
The distilled model retains broadly comparable appearance quality, but its effects differ by dataset. On Tanks and Temples, FID improves from 51.33 to 49.71 and LPIPS from 0.552 to 0.545, while Camera Controllability falls from 63.87 to 58.12 and Style Consistency from 85.07 to 78.91.
- On DL3DV, the distilled model has FID 43.63 versus 43.43 and LPIPS 0.507 versus 0.498, but improves Subjective Quality from 44.54 to 45.21 and Style Consistency from 87.46 to 88.57.
- Distillation also lowers SSIM and increases reprojection error on both datasets. Its speed advantage therefore comes with measurable trade-offs rather than uniform preservation of the full model’s performance.
5.3. Evaluation on 3D Scene Generation
3D scene evaluation pairs video-generation baselines with DAv3 and measures novel views rendered from the reconstructed 3DGS. LPIPS-G compares rendered views against ground-truth frames, whereas LPIPS-P compares them against generated video frames to measure how faithfully those observations can be reconstructed.
- FID and Subjective Quality measure rendered appearance.
- LPIPS-P is a reconstruction-based proxy for video consistency, not a direct measurement of ground-truth geometry; it also depends on the reconstruction model.
- Lyra and FantasyWorld are additionally compared qualitatively because their short-video pipelines limit scene scale.
Table 2 reports the best values for Ours Full across all listed reconstruction metrics on both datasets. Compared with Ours + DAv3, fine-tuning improves DL3DV LPIPS-P from 0.413 to 0.381, LPIPS-G from 0.603 to 0.579, FID from 74.39 to 65.94, and Subjective Quality from 17.02 to 20.52.
- On Tanks and Temples, the same comparison improves LPIPS-P from 0.409 to 0.372, LPIPS-G from 0.648 to 0.629, FID from 79.36 to 72.47, and Subjective Quality from 14.42 to 18.80.
- Ours + DAv3 already outperforms the other listed video-plus-DAv3 pipelines, separating the benefit of improved input videos from the additional reconstruction adaptation.
- Figure 4 shows cleaner building and locomotive renderings with Ours Full. Figure 5 illustrates larger spatial coverage than Lyra and FantasyWorld, but does not provide a quantitative scene-scale benchmark.
Using the same pretrained DAv3 reconstruction stage, Lyra 2.0 videos yield better scores than the listed competing video pipelines. Ours Full further improves every listed metric on both datasets, separating the benefit of reconstruction adaptation from gains due to video generation.
Building and locomotive renderings show cleaner structure with Lyra 2.0 video inputs under the shared DAv3 pipeline. The fine-tuned Ours Full variant further sharpens structure and reduces visible artifacts, complementing its quantitative comparison with Ours + DAv3.
Red boxes identify approximately the same starting region across methods, making Lyra 2.0’s additional spatial coverage visible. FantasyWorld is shown as a point cloud, whereas Lyra and Lyra 2.0 use 3DGS, so the comparison primarily illustrates scene extent rather than controlled rendering quality.
5.4. Ablation Study
Ablations on Tanks and Temples support the combined use of per-frame memory, learned correspondence aggregation, FramePack, and self-augmentation. Table 3 shows that the full model leads on SSIM, Style Consistency, and Camera Controllability, while ablated variants lead on other metrics.
- w/ Global Point Cloud replaces both the per-frame cache and correspondence conditioning with a fused cloud and rendered images. Camera Controllability drops from 63.87 to 49.86 and Style Consistency from 85.07 to 82.42, although reprojection error improves from 0.069 to 0.067.
- w/ Explicit Corr. Fusion replaces learned aggregation with hard depth-based correspondence fusion. Camera Controllability falls to 57.29, while FID improves to 49.13 from 51.33.
- w/o FramePack lowers Style Consistency to 80.61 and increases reprojection error to 0.079.
- w/o Self-Augmentation lowers Style Consistency to 77.98 and Camera Controllability to 53.92, despite improving Subjective Quality to 47.88 and reprojection error to 0.066. Figure 6 illustrates pose errors and appearance drift in the ablated variants.
The full model has the strongest SSIM, Style Consistency, and Camera Controllability, but ablated variants lead on other metrics. Removing self-augmentation raises Subjective Quality to 47.88 and lowers reprojection error to 0.066 while reducing Style Consistency to 77.98, showing that individual-frame quality and long-horizon stability need separate evaluation.
The building and horse-statue examples reveal viewpoint and appearance deviations in the ablated designs. They provide qualitative context for reduced camera-control and style-consistency scores, without establishing that every ablation is worse on every metric.
5.5. Applications
The application demonstrations combine trajectory planning, progressive scene expansion, and export to NVIDIA Isaac Sim. The GUI visualizes accumulated point clouds so users can select revisits or unexplored regions; Figure 7 follows the pipeline from interactive generation through surface-mesh extraction to simulation.
- Figures 1 and 8 show indoor, outdoor, and imaginative inputs beyond the evaluation benchmarks, including multiple trajectories departing from the same initial image.
- Exported 3DGS and mesh assets support the demonstrated robot simulation, but the paper does not report downstream policy-training gains or a quantitative simulation-validity evaluation.
- Interactive exploration is segment-based: the distilled model takes approximately 15 seconds to produce an 80-frame segment on the reported GPU.
The kitchen example connects trajectory selection in the GUI to an extracted surface mesh and an Isaac Sim scene. It demonstrates an asset-export workflow for embodied simulation, but does not quantify collision accuracy, physical properties, or downstream robot-policy performance.
The illustrated environments extend beyond the real-world benchmark scenes and include two trajectories from a common starting image in the second example. Their combined reconstruction demonstrates branching exploration qualitatively, without a numerical out-of-distribution consistency evaluation.
6. Discussion
The discussion identifies two remaining limitations: the framework targets static environments, and it can reproduce exposure variations present in DL3DV. These photometric inconsistencies can create artifacts when generated views are lifted into 3DGS.
- The authors suggest network-level photometric stabilization or photometrically consistent synthetic training data as possible remedies, rather than presenting evaluated solutions.
- Dynamic scenes are left for future work; the demonstrated reconstruction and simulator export do not constitute dynamic world modeling.
Appendix
- A. Implementation Details — A.1. Model Architecture: The backbone is Wan 2.1-14B, and all training and inference use 832×480 pixels. Depth-warped recent-frame images are VAE-encoded and concatenated with the denoising latent. Plücker rays and canonical coordinate maps are embedded using pixel-shuffle and linear operations, with sinusoidal positional encoding for coordinate-map channels; their embeddings enter attention queries and keys at every transformer block, not values. Figure 9 evaluates retrieval coverage using the latter half of training videos as targets and ground-truth depth for coverage checks, supporting N_s = 5 as a coverage–efficiency compromise.
- A.2. Training: Spatial memory uses N_s = 5, d = 8, and δ = 0.1 in normalized depth units; self-augmentation uses p_aug = 0.7. AdamW uses learning rate 3×10−5 and weight decay 0.1, with batch size 64 across 64 NVIDIA GB200 GPUs for 7,000 iterations. Newly added modules are zero-initialized, and training uses bf16 mixed precision. Rectified flow matching samples timesteps from a logit-normal distribution with mean 0 and standard deviation 1 in logit space, with uniform time weighting.
- A.3. Inference: The full model uses FlowUniPC with 35 steps and text classifier-free guidance scale 5.0. One 80-frame segment takes approximately 194 seconds on a single NVIDIA GB200 GPU, including depth estimation, retrieval, and denoising; Ours DMD takes approximately 15 seconds with 4 steps and no separate CFG passes. Spatial-memory retrieval takes less than 1 second per segment in both cases.
- A.4. 3D Reconstruction: The Gaussian head uses k = 2 for a 4× reduction in Gaussian count. DAv3 fine-tuning uses 3,000 autoregressively generated one-minute videos based on DL3DV images and camera trajectories, with 10,000 iterations, learning rate 5×10−5, and batch size 8. Mesh hierarchy levels and voxel sizes depend on scene scale, and marching-cubes surfaces are merged at transitions between levels.
- A.5. Related Work: The extended review traces 3D generation from category-specific GANs through CLIP supervision, Score Distillation Sampling, multi-view diffusion, iterative inpainting, and explicit NeRF, Gaussian, or mesh representations. It also reviews feed-forward 3D models, noting their frequent restriction to static scenes and the difficulty of generalizing dynamic reconstruction to diverse generated content and large viewpoint changes.
Brief Thoughts
The main contribution is the separation of spatial-memory geometry from appearance synthesis, paired with inexpensive training on imperfect histories. Video comparisons and ablations support improved appearance and long-range style consistency, while reconstruction comparisons show consistent benefits from adapting DAv3. Camera accuracy, reprojection error, and individual-frame quality remain distinct objectives with measurable trade-offs. The global-cloud ablation changes both memory representation and conditioning, so it supports the combined design rather than isolating the effect of per-frame storage alone. The demonstrated result is interactive, segment-based creation of static 3D assets. Finite-horizon comparisons and qualitative simulator demonstrations do not establish unlimited persistence or physical simulation fidelity.