用三维表面点索引记忆,让视频生成更连贯、高效。
VMem: Consistent Interactive Video Scene Generation with Surfel-Indexed View Memory
- 用表面点(surfels)几何索引过往视角,精准存储与检索。
- 长时视频生成中保持场景一致性,计算成本降低显著。
- 适合需要长期连贯性的交互式场景生成任务。
我们提出一种新型记忆模块,用于构建可交互探索环境的视频生成模型。以往方法要么通过逐步重建3D结构来外推2D视图,但会快速积累误差;要么使用短上下文窗口的视频生成器,难以维持长期场景连贯性。为解决这些问题,我们引入基于3D表面元素(surfels)几何索引的视图记忆(VMem),将过去视角按其观测到的表面点进行索引存储。该机制能高效检索生成新视图时最相关的过往视角。仅聚焦于这些相关视图,我们的方法在极低计算成本下实现对想象场景的一致性探索。我们在具有挑战性的长期场景合成基准上评估,结果表明,在维持场景连贯性和相机控制方面优于现有方法。
原文摘要 · Abstract (English)
We propose a novel memory module for building video generators capable of interactively exploring environments. Previous approaches have achieved similar results either by out-painting 2D views of a scene while incrementally reconstructing its 3D geometry-which quickly accumulates errors-or by using video generators with a short context window, which struggle to maintain scene coherence over the long term. To address these limitations, we introduce Surfel-Indexed View Memory (VMem), a memory module that remembers past views by indexing them geometrically based on the 3D surface elements (surfels) they have observed. VMem enables efficient retrieval of the most relevant past views when generating new ones. By focusing only on these relevant views, our method produces consistent explorations of imagined environments at a fraction of the computational cost required to use all past views as context. We evaluate our approach on challenging long-term scene synthesis benchmarks and demonstrate superior performance compared to existing methods in maintaining scene coherence and camera control.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。