通过测试时训练实现视频流中空间信息的持续更新与组织。
Spatial-TTT: Streaming Visual-based Spatial Intelligence with Test-Time Training
- 用快速权重和滑动窗口注意力动态捕捉长期视觉空间证据。
- 在多个视频空间基准上达到当前最优性能,显著提升长时序理解能力。
- 适合研究视频理解、空间智能与测试时训练的学者与工程师。
人类通过连续的视觉观察感知和理解真实空间,因此从可能无限的视频流中持续维护和更新空间证据,是实现空间智能的关键。核心挑战不仅在于更长的上下文窗口,更在于如何选择、组织和保留空间信息。本文提出 Spatial-TTT,一种基于测试时训练(TTT)的流式视觉空间智能方法,通过调整部分参数(快速权重)来捕捉并组织长时程场景视频中的空间证据。我们设计了混合架构,结合大块更新与滑动窗口注意力,实现高效的视频处理。为进一步增强空间感知,引入3D时空卷积的时空预测机制,促进模型捕捉帧间的几何对应与时间连续性。此外,我们构建了一个包含密集3D空间描述的数据集,引导模型以结构化方式记忆和组织全局3D空间信号。大量实验表明,Spatial-TTT显著提升了长时序空间理解能力,在多个视频空间基准上达到领先水平。
原文摘要 · Abstract (English)
Humans perceive and understand real-world spaces through a stream of visual observations. Therefore, the ability to streamingly maintain and update spatial evidence from potentially unbounded video streams is essential for spatial intelligence. The core challenge is not simply longer context windows but how spatial information is selected, organized, and retained over time. In this paper, we propose Spatial-TTT towards streaming visual-based spatial intelligence with test-time training (TTT), which adapts a subset of parameters (fast weights) to capture and organize spatial evidence over long-horizon scene videos. Specifically, we design a hybrid architecture and adopt large-chunk updates parallel with sliding-window attention for efficient spatial video processing. To further promote spatial awareness, we introduce a spatial-predictive mechanism applied to TTT layers with 3D spatiotemporal convolution, which encourages the model to capture geometric correspondence and temporal continuity across frames. Beyond architecture design, we construct a dataset with dense 3D spatial descriptions, which guides the model to update its fast weights to memorize and organize global 3D spatial signals in a structured manner. Extensive experiments demonstrate that Spatial-TTT improves long-horizon spatial understanding and achieves state-of-the-art performance on video spatial benchmarks. Project page: https://liuff19.github.io/Spatial-TTT.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。