用注意力机制实现灵活视频记忆检索,提升长时序建模效果
Compression and Retrieval: Implicit Memory Retrieval for Video World Models

- 通过位置编码注入视角信息,用注意力动态检索记忆
- 压缩网络降低计算开销,支持长序列处理
- 在真实相机轨迹数据集上表现优异,通用性强
视频世界模型有望模拟交互式环境,但维持复杂相机轨迹下的长期一致性记忆仍是关键挑战。现有方法通常依赖计算成本高昂的上下文扩展或僵化的启发式检索机制,难以泛化到不同相机轨迹和场景。本文提出压缩与检索(CaR)机制,一种基于注意力的隐式记忆检索方法。通过位置编码注入视角信息,实现灵活的记忆检索。为高效处理长上下文且计算开销极小,进一步引入轻量级上下文压缩网络。此外,构建了包含真实相机轨迹和帧级标注的大规模合成数据集SceneFly,用于训练与评估长时序视频世界模型。大量实验表明,该方法在主流基准上达到当前最优性能,并展现出对开放域场景的强大泛化能力。
原文摘要 · Abstract (English)
Video world models hold promise for simulating interactive environments, yet maintaining consistent long-term memory across complex camera trajectories remains a critical challenge. Existing methods typically rely on computationally expensive context scaling or rigid heuristic retrieval mechanisms, which lacks generalization to varying camera trajectories and environments. In this paper, we propose Compression and Retrieval (CaR), an attention-driven implicit memory retrieval mechanism to overcome these limitations. By injecting viewpoint information via positional encoding, our method performs flexible memory retrieval through attention computation. To efficiently process extended contexts with minimal computational overhead, we further introduce a lightweight context compression network. Furthermore, we construct SceneFly, a large-scale synthetic dataset featuring realistic camera trajectories and frame-level annotations to train and evaluate long-horizon video world models. Extensive experiments demonstrate that our approach achieves state-of-the-art results on established benchmarks and exhibits strong generalization to open-domain scenes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。