用可演化3D记忆生成连贯全景视频,实现长期空间一致的虚拟探索。
EvoWorld: Evolving Panoramic World Generation with Explicit 3D Memory
- 引入显式3D记忆,通过几何重投影提供空间引导。
- 在合成场景、室内和真实环境上均显著提升视觉真实感与空间一致性。
- 适合做长时序3D世界建模、虚拟现实或机器人导航的研究者。
人类具备在脑中回放和探索曾经历的3D环境的能力。受此启发,我们提出EvoWorld:一种将全景视频生成与可演化3D记忆相结合的世界模型,实现空间一致的长时程探索。输入单张全景图像后,EvoWorld首先利用具备细粒度视角控制的视频生成器生成未来帧,再通过前馈式即插即用的Transformer演化场景的3D重建,并最终基于该动态3D记忆的几何重投影条件生成未来画面。不同于仅生成视频的现有方法,关键创新在于将演化中的3D重建作为显式空间引导,将重构几何投影至目标视角,提供丰富空间线索,显著提升视觉真实感与几何一致性。为评估长程探索能力,我们构建首个涵盖合成户外环境、Habitat室内场景及挑战性真实场景的综合性基准,重点考察环路闭合检测与长轨迹空间连贯性。大量实验表明,相比现有方法,我们的演化3D记忆显著提升视觉保真度并维持空间场景一致性,代表了长时程空间一致世界建模的重要进展。
原文摘要 · Abstract (English)
Humans possess a remarkable ability to mentally explore and replay 3D environments they have previously experienced. Inspired by this mental process, we present EvoWorld: a world model that bridges panoramic video generation with evolving 3D memory to enable spatially consistent long-horizon exploration. Given a single panoramic image as input, EvoWorld first generates future video frames by leveraging a video generator with fine-grained view control, then evolves the scene's 3D reconstruction using a feedforward plug-and-play transformer, and finally synthesizes futures by conditioning on geometric reprojections from this evolving explicit 3D memory. Unlike prior state-of-the-arts that synthesize videos only, our key insight lies in exploiting this evolving 3D reconstruction as explicit spatial guidance for the video generation process, projecting the reconstructed geometry onto target viewpoints to provide rich spatial cues that significantly enhance both visual realism and geometric consistency. To evaluate long-range exploration capabilities, we introduce the first comprehensive benchmark spanning synthetic outdoor environments, Habitat indoor scenes, and challenging real-world scenarios, with particular emphasis on loop-closure detection and spatial coherence over extended trajectories. Extensive experiments demonstrate that our evolving 3D memory substantially improves visual fidelity and maintains spatial scene coherence compared to existing approaches, representing a significant advance toward long-horizon spatially consistent world modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。