arXiv:2409.02371cs.CV2024-09被引 5

通过泰勒展开学习视频动态,提升自监督表示能力

Unfolding Videos Dynamics via Taylor Expansion

  • 用帧序列的高阶时间导数构建泰勒展开,捕捉运动特征
  • 在UCF101和Kinetics上显著提升动作识别与检测性能
  • 无需大模型或海量数据,适配现有自监督框架

受物理运动启发,我们提出一种新的自监督视频动态学习策略:视频时间微分实例判别(ViDiDi)。ViDiDi是一种简单且数据高效的策略,可直接应用于基于实例判别的现有自监督视频表征学习框架。其核心思想是通过视频帧序列的不同阶时间导数观察视频的多个方面。这些导数与原始帧共同支持离散时间下潜在连续动态的泰勒级数展开,其中高阶导数强调高阶运动特征。ViDiDi训练一个单一神经网络,采用平衡交替学习算法,将视频及其时间导数编码为一致嵌入。通过学习原始帧与导数的一致表示,编码器被引导关注运动特征而非静态背景,并揭示原始帧中的隐藏动态。因此,视频表示更依赖于动态特征进行区分。我们将ViDiDi集成到现有实例判别框架(VICReg、BYOL、SimCLR)中,在UCF101或Kinetics上进行预训练,并在视频检索、动作识别和动作检测等标准基准上测试。结果表明,性能显著提升,且无需使用大模型或大量数据。

原文摘要 · Abstract (English)

Taking inspiration from physical motion, we present a new self-supervised dynamics learning strategy for videos: Video Time-Differentiation for Instance Discrimination (ViDiDi). ViDiDi is a simple and data-efficient strategy, readily applicable to existing self-supervised video representation learning frameworks based on instance discrimination. At its core, ViDiDi observes different aspects of a video through various orders of temporal derivatives of its frame sequence. These derivatives, along with the original frames, support the Taylor series expansion of the underlying continuous dynamics at discrete times, where higher-order derivatives emphasize higher-order motion features. ViDiDi learns a single neural network that encodes a video and its temporal derivatives into consistent embeddings following a balanced alternating learning algorithm. By learning consistent representations for original frames and derivatives, the encoder is steered to emphasize motion features over static backgrounds and uncover the hidden dynamics in original frames. Hence, video representations are better separated by dynamic features. We integrate ViDiDi into existing instance discrimination frameworks (VICReg, BYOL, and SimCLR) for pretraining on UCF101 or Kinetics and test on standard benchmarks including video retrieval, action recognition, and action detection. The performances are enhanced by a significant margin without the need for large models or extensive datasets.

视频表征自监督学习动态建模泰勒展开

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。