arXiv:2603.17117cs.CV2026-03被引 14

混合空间记忆让视频世界模型更准地跟踪移动物体和相机运动。

MosaicMem: Hybrid Spatial Memory for Controllable Video World Models

  • 用分块拼接方式将图像块提升到3D,实现精准定位与检索。
  • 在相机运动下保持姿态一致,动态建模能力优于传统方法。
  • 适合需要长时一致性、可控编辑的视频生成任务。

视频扩散模型正从短时合理片段迈向需在相机运动、回访和干预下保持一致性的世界模拟器。然而空间记忆仍是关键瓶颈:显式3D结构虽能提升重投影一致性,却难以刻画运动物体;隐式记忆即使姿态正确,也常导致相机运动不准。我们提出混合空间记忆 MosaicMem,将图像块提升至3D以实现可靠定位与定向检索,同时利用模型原生条件控制保持提示遵循生成。MosaicMem通过分块拼接接口,在查询视角中组合空间对齐的图像块,保留应持续的内容,允许模型补全应变化的部分。结合PRoPE相机条件及两种新记忆对齐方法,实验显示其姿态遵循性优于隐式记忆,动态建模强于显式基线。MosaicMem还支持分钟级导航、基于记忆的场景编辑与自回归滚动生成。

原文摘要 · Abstract (English)

Video diffusion models are moving beyond short, plausible clips toward world simulators that must remain consistent under camera motion, revisits, and intervention. Yet spatial memory remains a key bottleneck: explicit 3D structures can improve reprojection-based consistency but struggle to depict moving objects, while implicit memory often produces inaccurate camera motion even with correct poses. We propose Mosaic Memory (MosaicMem), a hybrid spatial memory that lifts patches into 3D for reliable localization and targeted retrieval, while exploiting the model's native conditioning to preserve prompt-following generation. MosaicMem composes spatially aligned patches in the queried view via a patch-and-compose interface, preserving what should persist while allowing the model to inpaint what should evolve. With PRoPE camera conditioning and two new memory alignment methods, experiments show improved pose adherence compared to implicit memory and stronger dynamic modeling than explicit baselines. MosaicMem further enables minute-level navigation, memory-based scene editing, and autoregressive rollout.

视频生成世界模型空间记忆扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。