用局部几何记忆解决长视频生成中的空间不一致问题
AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- 通过检索与轨迹对齐的局部几何记忆替代全局重建
- 多锚点编织控制器融合记忆,提升长期一致性
- 适合需要高空间一致性视频生成的研究者
长时序相机可控视频生成中保持空间世界一致性仍是核心挑战。现有基于记忆的方法通常依赖历史视图重建的全局3D场景渲染锚定视频进行条件生成,但多视角重建不可避免引入跨视角错位:姿态与深度估计误差导致同一表面在不同视角被重建在略微不同的3D位置。这些不一致在融合后累积为噪声几何,污染条件信号并降低生成质量。我们提出AnchorWeave,一种记忆增强型视频生成框架,将单一存在错位的全局记忆替换为多个干净的局部几何记忆,并学习协调其跨视角不一致。为此,AnchorWeave执行与目标轨迹对齐的覆盖驱动式局部记忆检索,并在生成过程中通过多锚点编织控制器整合所选局部记忆。大量实验表明,AnchorWeave显著提升长期场景一致性,同时保持强视觉质量;消融与分析研究进一步验证了局部几何条件、多锚点控制及覆盖驱动检索的有效性。
原文摘要 · Abstract (English)
Maintaining spatial world consistency over long horizons remains a central challenge for camera-controllable video generation. Existing memory-based approaches often condition generation on globally reconstructed 3D scenes by rendering anchor videos from the reconstructed geometry in the history. However, reconstructing a global 3D scene from multiple views inevitably introduces cross-view misalignment, as pose and depth estimation errors cause the same surfaces to be reconstructed at slightly different 3D locations across views. When fused, these inconsistencies accumulate into noisy geometry that contaminates the conditioning signals and degrades generation quality. We introduce AnchorWeave, a memory-augmented video generation framework that replaces a single misaligned global memory with multiple clean local geometric memories and learns to reconcile their cross-view inconsistencies. To this end, AnchorWeave performs coverage-driven local memory retrieval aligned with the target trajectory and integrates the selected local memories through a multi-anchor weaving controller during generation. Extensive experiments demonstrate that AnchorWeave significantly improves long-term scene consistency while maintaining strong visual quality, with ablation and analysis studies further validating the effectiveness of local geometric conditioning, multi-anchor control, and coverage-driven retrieval.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。