提出快速空间记忆模型,实现长序列3D/4D重建的高效稳定适应。
Fast Spatial Memory with Elastic Test-Time Training
- 引入弹性测试时训练,用锚点状态与费雪加权先验稳定快速权重更新。
- 支持小块输入下长序列重建,显著缓解灾难性遗忘与相机插值捷径问题。
- 适合需要实时3D/4D场景建模的机器人、AR/VR应用,兼具高效与鲁棒性。
大块测试时训练(LaCT)在长上下文3D重建中表现优异,但其全可塑的推理时更新易导致灾难性遗忘与过拟合。现有方法通常仅使用单一大块覆盖全输入序列,难以实现任意长序列的单次处理。本文受弹性权重固化启发,提出弹性测试时训练,通过费雪加权的弹性先验约束快速权重更新,并以指数移动平均维护锚点状态,在稳定性与可塑性间取得平衡。基于此架构,我们提出快速空间记忆(FSM),一种高效可扩展的4D重建模型,能从长观测序列中学习时空表示并渲染新视角-时间组合。我们在大规模精心筛选的3D/4D数据上预训练了FSM,以捕捉复杂空间环境的动力学与语义。大量实验表明,FSM可在小块输入下实现快速适应,提供高质量3D/4D重建,有效缓解相机插值捷径问题。本工作推动LaCT从受限单块设置迈向稳健的多块适应,是实现真正长序列泛化的重要一步,同时显著缓解激活-内存瓶颈。
原文摘要 · Abstract (English)
Large Chunk Test-Time Training (LaCT) has shown strong performance on long-context 3D reconstruction, but its fully plastic inference-time updates remain vulnerable to catastrophic forgetting and overfitting. As a result, LaCT is typically instantiated with a single large chunk spanning the full input sequence, falling short of the broader goal of handling arbitrarily long sequences in a single pass. We propose Elastic Test-Time Training inspired by elastic weight consolidation, that stabilizes LaCT fast-weight updates with a Fisher-weighted elastic prior around a maintained anchor state. The anchor evolves as an exponential moving average of past fast weights to balance stability and plasticity. Based on this updated architecture, we introduce Fast Spatial Memory (FSM), an efficient and scalable model for 4D reconstruction that learns spatiotemporal representations from long observation sequences and renders novel view-time combinations. We pre-trained FSM on large-scale curated 3D/4D data to capture the dynamics and semantics of complex spatial environments. Extensive experiments show that FSM supports fast adaptation over long sequences and delivers high-quality 3D/4D reconstruction with smaller chunks and mitigating the camera-interpolation shortcut. Overall, we hope to advance LaCT beyond the bounded single-chunk setting toward robust multi-chunk adaptation, a necessary step for generalization to genuinely longer sequences, while substantially alleviating the activation-memory bottleneck.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。