arXiv:2603.23413cs.CV2026-03

无需显式建模3D,用隐式记忆实现视频场景一致生成

I3DM: Implicit 3D-aware Memory Retrieval and Injection for Consistent Video Scene Generation

  • 利用预训练模型中间特征评分视角相关性,实现强遮挡下的稳定记忆检索
  • 通过隐式投影历史帧到目标视角,提升重访一致性与相机控制精度
  • 适合需要长时序一致性的视频生成任务,如虚拟拍摄与数字人

尽管视频生成技术取得显著进展,但在重复访问已探索区域时保持长期场景一致性仍具挑战。现有方法要么依赖显式构建3D几何结构,存在误差累积和尺度模糊问题;要么采用简单的相机视场(FoV)检索,在复杂遮挡下表现不佳。为此,我们提出I3DM,一种新型的隐式3D感知记忆机制,用于一致视频场景生成,绕过显式3D重建。核心是基于预训练前馈新视角合成(FF-NVS)模型中间特征的3D感知记忆检索策略,可有效评估视角相关性,即使在高度遮挡场景下仍能稳健检索。此外,为充分使用检索到的历史帧,我们引入3D对齐的记忆注入模块,该模块隐式将历史内容映射至目标视角,并自适应地在可靠映射区域上条件化生成,显著提升重访一致性与相机控制精度。大量实验表明,本方法优于当前最先进方法,在重访一致性、生成保真度和相机控制精度方面均取得更优表现。

原文摘要 · Abstract (English)

Despite remarkable progress in video generation, maintaining long-term scene consistency upon revisiting previously explored areas remains challenging. Existing solutions rely either on explicitly constructing 3D geometry, which suffers from error accumulation and scale ambiguity, or on naive camera Field-of-View (FoV) retrieval, which typically fails under complex occlusions. To overcome these limitations, we propose I3DM, a novel implicit 3D-aware memory mechanism for consistent video scene generation that bypasses explicit 3D reconstruction. At the core of our approach is a 3D-aware memory retrieval strategy, which leverages the intermediate features of a pre-trained Feed-Forward Novel View Synthesis (FF-NVS) model to score view relevance, enabling robust retrieval even in highly occluded scenarios. Furthermore, to fully utilize the retrieved historical frames, we introduce a 3D-aligned memory injection module. This module implicitly warps historical content to the target view and adaptively conditions the generation on reliable warping regions, leading to improved revisit consistency and accurate camera control. Extensive experiments demonstrate that our method outperforms state-of-the-art approaches, achieving superior revisit consistency, generation fidelity, and camera control precision.

视频生成3D感知记忆机制一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。