arXiv:2605.23045cs.CVcs.AI2026-05

用运动信息训练视频模型,数据量少1万倍仍达顶尖水平

The TIME Machine: On The Power of Motion for Efficient Perception

论文配图:The TIME Machine: On The Power of Motion for Efficient Perception
图 1 · 摘自论文原文
  • 以点轨迹为输入,通过掩码自编码器学习运动表征
  • 仅用合成数据训练,零样本测试性能媲美大模型
  • 摆脱语言依赖,适合需要细粒度时序理解的场景

视频表征学习近年进展显著,但受限于训练规模和语言依赖性。本文提出以运动为核心模态的新方法:利用视频中的点轨迹,通过掩码自编码器重建缺失轨迹,实现自监督学习。该方法不依赖外观信息,大幅降低数据需求;同时摆脱语言约束,提升对细粒度时序概念的建模能力。所提出的TIME(Temporally Informed Motion Embedding)嵌入仅在合成运动数据上训练,零样本测试表现与使用高达4个数量级更多数据的先进模型相当,是迈向更高效、更具时序感知能力视频模型的重要一步。

原文摘要 · Abstract (English)

Video representation learning has seen tremendous progress in recent years. This has been driven by many factors, including the scale of training and the success of visual models trained contrastively with language. While these factors have pushed the boundaries of what video models can do, they also introduce their own set of limitations: first, scaling video models can reach prohibitive costs and second, learning from language restricts the range of concepts that can be learned to those in captions. As a result, video models still struggle with temporal understanding. In this paper we propose a novel approach that uses motion as the central modality for video representation. In particular, given the motion in a video in the form of point-tracks, we use a masked-autoencoder to mask some of the tracks and train the autoencoder to reconstruct the missing tracks. This allows us to learn a representation in a self-supervised manner. We show that using motion to represent videos actually addresses both of the core limitations of video technology. First, it allows us to massively reduce the scale of training data, as motion is inherently appearance-independent and hence needs fewer examples to generalize well. Second, motion allows us to bypass the language-dependent training paradigm, learning better fine-grained concepts. The result is an embedding that we call TIME (Temporally Informed Motion Embedding), a representation trained exclusively on synthetic motion data. We test this embedding on a wide set of tasks in a zero-shot manner. We observe that without bells and whistles, performance is on par with state-of-the-art models using up to 4 orders of magnitude less training data. This is a stepping stone towards a new paradigm of video models that are both more temporally aware as well as more scalable.

视频表征自监督学习运动建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。