用视频生成3D场景,实现自由视角探索
WorldExplorer: Towards Generating Fully Navigable 3D Scenes
- 基于自回归视频轨迹生成,逐步构建可导航3D场景
- 支持大范围相机运动,视觉质量稳定无畸变
- 适合虚拟现实、游戏开发等需沉浸式场景的领域
从文本生成3D世界是计算机视觉的重要目标。现有方法在场景内探索时受限,超出中心或全景视角后会产生拉伸和噪声。为此,我们提出WorldExplorer,一种基于自回归视频轨迹生成的新方法,可构建全可导航、多视角视觉一致的3D场景。首先生成360度全景对应的多视图一致图像作为初始场景;随后通过迭代生成沿预设短轨迹的多段视频,深入探索场景并包含物体周围运动。新提出的场景记忆机制使每段视频基于最相关的历史视图,碰撞检测机制避免穿入物体等异常结果。最终通过3D高斯泼溅优化融合所有生成视图,形成统一3D表示。相比已有方法,WorldExplorer在大幅相机运动下仍保持高质量与稳定性,首次实现真实且无限制的场景探索。我们认为这标志着向生成沉浸式、真正可探索虚拟3D环境迈出关键一步。
原文摘要 · Abstract (English)
Generating 3D worlds from text is a highly anticipated goal in computer vision. Existing works are limited by the degree of exploration they allow inside of a scene, i.e., produce streched-out and noisy artifacts when moving beyond central or panoramic perspectives. To this end, we propose WorldExplorer, a novel method based on autoregressive video trajectory generation, which builds fully navigable 3D scenes with consistent visual quality across a wide range of viewpoints. We initialize our scenes by creating multi-view consistent images corresponding to a 360 degree panorama. Then, we expand it by leveraging video diffusion models in an iterative scene generation pipeline. Concretely, we generate multiple videos along short, pre-defined trajectories, that explore the scene in depth, including motion around objects. Our novel scene memory conditions each video on the most relevant prior views, while a collision-detection mechanism prevents degenerate results, like moving into objects. Finally, we fuse all generated views into a unified 3D representation via 3D Gaussian Splatting optimization. Compared to prior approaches, WorldExplorer produces high-quality scenes that remain stable under large camera motion, enabling for the first time realistic and unrestricted exploration. We believe this marks a significant step toward generating immersive and truly explorable virtual 3D environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。