arXiv:2502.07811cs.CV2025-02被引 2

用视频与帧图像联合训练,让模型更好理解动作的语义细节。

CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders

  • 跨模态对比学习,融合视频时序与帧图像空间信息
  • 在多个数据集上超越现有最佳方法,提升动作理解能力
  • 适合需要精细动作识别的研究者和开发者

当前基于视频的掩码自编码器(MAE)主要从视觉角度学习时空表征,易忽略动作特有的语义属性,如特定交互或序列,难以捕捉上下文丰富且连续的动作本质。人类能将静态图像中的视觉概念、物体视角不变性及语义属性映射到动态场景中。现有视频与图像MAE使用独立数据集,缺乏充分语义支持。为此,我们提出CrossVideoMAE,一种端到端自监督跨模态对比学习的MAE,可同时学习视频级与帧级丰富的时空表征与语义属性。通过在特征不变空间中整合视频的时空信息与采样帧的空间信息,并增强视频域内变换的不变性,实现可见标记项的联合嵌入与模态内/间特征对应,从而在无标签条件下从视频与帧图像中获取强引导信号。大量实验表明,该方法优于现有最先进方法,消融实验验证了其有效性。

原文摘要 · Abstract (English)

Current video-based Masked Autoencoders (MAEs) primarily focus on learning effective spatiotemporal representations from a visual perspective, which may lead the model to prioritize general spatial-temporal patterns but often overlook nuanced semantic attributes like specific interactions or sequences that define actions - such as action-specific features that align more closely with human cognition for space-time correspondence. This can limit the model's ability to capture the essence of certain actions that are contextually rich and continuous. Humans are capable of mapping visual concepts, object view invariance, and semantic attributes available in static instances to comprehend natural dynamic scenes or videos. Existing MAEs for videos and static images rely on separate datasets for videos and images, which may lack the rich semantic attributes necessary for fully understanding the learned concepts, especially when compared to using video and corresponding sampled frame images together. To this end, we propose CrossVideoMAE an end-to-end self-supervised cross-modal contrastive learning MAE that effectively learns both video-level and frame-level rich spatiotemporal representations and semantic attributes. Our method integrates mutual spatiotemporal information from videos with spatial information from sampled frames within a feature-invariant space, while encouraging invariance to augmentations within the video domain. This objective is achieved through jointly embedding features of visible tokens and combining feature correspondence within and across modalities, which is critical for acquiring rich, label-free guiding signals from both video and frame image modalities in a self-supervised manner. Extensive experiments demonstrate that our approach surpasses previous state-of-the-art methods and ablation studies validate the effectiveness of our approach.

自监督学习视频理解跨模态表征学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。