arXiv:2506.04225cs.CV2025-06被引 100

用单图生成可自由探索的3D动态场景,全程保持空间一致性。

Voyager: Long-Range and World-Consistent Video Diffusion for Explorable 3D Scene Generation

  • 联合生成彩色与深度视频,保证帧间世界一致
  • 支持长距离探索,自动扩展场景并保持上下文连贯
  • 无需人工标注,可大规模训练提升生成质量

现实应用如游戏和虚拟现实常需用户沿自定义相机轨迹探索3D场景。尽管已有方法能从文本或图像生成3D物体,但生成长距离、几何一致且可交互的3D场景仍是难题。本文提出Voyager,一种新型视频扩散框架,仅凭一张图像即可生成用户定义相机路径下的世界一致3D点云序列。不同于现有方法,Voyager实现端到端生成与重建,无需依赖结构光或多视图立体等3D重建流程。其核心包括:1)世界一致视频扩散:统一架构联合生成对齐的RGB与深度视频序列,基于已有世界观测确保全局连贯性;2)长程世界探索:采用高效世界缓存结合点云剔除,以及上下文感知的自回归推理与平滑采样,实现迭代场景扩展;3)可扩展数据引擎:自动化相机位姿估计与度量深度预测,适用于任意视频,无需人工3D标注,支持大规模多样化训练数据构建。整体显著优于现有方法,在视觉质量和几何精度上均有提升,具备广泛应用潜力。

原文摘要 · Abstract (English)

Real-world applications like video gaming and virtual reality often demand the ability to model 3D scenes that users can explore along custom camera trajectories. While significant progress has been made in generating 3D objects from text or images, creating long-range, 3D-consistent, explorable 3D scenes remains a complex and challenging problem. In this work, we present Voyager, a novel video diffusion framework that generates world-consistent 3D point-cloud sequences from a single image with user-defined camera path. Unlike existing approaches, Voyager achieves end-to-end scene generation and reconstruction with inherent consistency across frames, eliminating the need for 3D reconstruction pipelines (e.g., structure-from-motion or multi-view stereo). Our method integrates three key components: 1) World-Consistent Video Diffusion: A unified architecture that jointly generates aligned RGB and depth video sequences, conditioned on existing world observation to ensure global coherence 2) Long-Range World Exploration: An efficient world cache with point culling and an auto-regressive inference with smooth video sampling for iterative scene extension with context-aware consistency, and 3) Scalable Data Engine: A video reconstruction pipeline that automates camera pose estimation and metric depth prediction for arbitrary videos, enabling large-scale, diverse training data curation without manual 3D annotations. Collectively, these designs result in a clear improvement over existing methods in visual quality and geometric accuracy, with versatile applications.

3D生成视频扩散世界一致性可探索场景

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。