arXiv:2602.14705cs.CV2026-02被引 2

用长时间运动信息提升视觉感知,效果优于图像

It's a Matter of Time: Three Lessons on Long-Term Motion for Perception

  • 利用长期运动轨迹学习视觉表征
  • 低数据和零样本任务中泛化能力更强
  • 运动信息维度低,效率高于传统视频模型

时间信息对感知至关重要,但其作用仍不明确:从长期运动中能学到什么?它在视觉学习中有哪些特性?我们借助点轨迹估计的进展,探索多种感知任务。得出三点启示:1)长期运动表征蕴含动作、物体、材质和空间信息,有时甚至优于图像;2)在低数据设置和零样本任务中,运动表征的泛化能力显著优于图像表征;3)运动信息维度极低,相比标准视频表征,在计算量(GFLOPs)与精度之间有更好权衡,二者结合性能优于视频表征单独使用。这些发现为未来感知模型设计提供了新方向。

原文摘要 · Abstract (English)

Temporal information has long been considered to be essential for perception. While there is extensive research on the role of image information for perceptual tasks, the role of the temporal dimension remains less well understood: What can we learn about the world from long-term motion information? What properties does long-term motion information have for visual learning? We leverage recent success in point-track estimation, which offers an excellent opportunity to learn temporal representations and experiment on a variety of perceptual tasks. We draw 3 clear lessons: 1) Long-term motion representations contain information to understand actions, but also objects, materials, and spatial information, often even better than images. 2) Long-term motion representations generalize far better than image representations in low-data settings and in zero-shot tasks. 3) The very low dimensionality of motion information makes motion representations a better trade-off between GFLOPs and accuracy than standard video representations, and used together they achieve higher performance than video representations alone. We hope these insights will pave the way for the design of future models that leverage the power of long-term motion information for perception.

视觉感知运动表征零样本学习高效模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。