提出在线记忆机制,实现实时动态视角合成与长期记忆保持。
Online Neural Space Time Memory for Dynamic Novel View Synthesis

- 周期性更新记忆,每帧应用记忆,降低计算负担
- 在动态人体场景中实现毫秒级在线记忆与实时合成
- 适合需要长期记忆的实时视频重建任务
从多视角流视频进行在线新视角合成面临根本矛盾:需维持持久、长时程记忆以重建临时遮挡区域,同时满足严格的实时性要求。尽管测试时训练(TTT)提供强大记忆机制,但标准模型要求每帧都进行基于梯度的记忆更新以适应动态场景变化,计算开销大且易引发长时间上下文下的不稳定性。鉴于记忆更新比记忆应用更耗资源,且视频内容高度冗余,我们提出解耦两者频率:仅周期性更新记忆,而每帧仍应用记忆,并使用跨视角注意力处理先验记忆状态与当前帧间的形变。为锁定历史上下文,引入辅助记忆损失以强制场景持续内化,并采用记忆缓存策略,防止活跃权重发生灾难性漂移。方法在含动态人体运动的场景中实现实时、业界领先性能,支持分钟级在线记忆能力。
原文摘要 · Abstract (English)
Online novel view synthesis from multi-view streaming videos faces a fundamental trade-off: maintaining a persistent, long-horizon memory to reconstruct temporarily occluded regions while operating under strict real-time constraints. While Test-Time Training (TTT) offers a powerful memory mechanism, standard models mandate gradient-based memory updates at every frame to adapt to the changing motion in dynamic scenes. The computational cost of heavy memory updates precludes real-time application and can lead to instability over long contexts. Given that memory updates are more demanding than memory application and video content is largely redundant, we propose to decouple the frequencies of these two processes. Our approach performs periodic memory updates while applying the memory on a per-frame basis, using cross-view attention to manage deformations between the prior memory state and the current frame. To lock in the historical context, we introduce two critical mechanisms: an auxiliary Memory Loss that forces persistent internalization of the scene, and a Memory Caching strategy that regularizes active weights against catastrophic drift. Our method demonstrates real-time, state-of-the-art performance on scenes with dynamic human motion as well as minute-scale online memorization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。