arXiv:2504.06153cs.CV2025-04CVPR被引 4

建立统一基准,系统分析视频自监督学习的关键影响因素。

A Large-Scale Analysis on Contextual Self-Supervised Video Representation Learning

  • 构建统一评估框架,公平比较六种方法与六类模型。
  • 发现数据量、噪声、分布对模型性能有显著影响。
  • 适用于研究自监督视频表征与大模型训练的学者。

自监督学习已成为无需标注的视频模型预训练强大范式,但现有方法实验设置各异,缺乏统一基准导致难以直接比较。本文建立统一基准,系统研究视频自监督学习的五个关键方面:数据集规模、模型复杂度、数据分布、数据噪声和特征表示。在五大数据集上,评估六种自监督方法与六种网络架构,涵盖两个下游任务。分析揭示预训练策略、数据特性、前置任务与模型结构间的交互关系。进一步扩展至视频基础模型(ViFMs),验证其在大规模视频表征学习中的适用性。基于这些洞察,提出新方法,在仅使用10%更少预训练数据的情况下超越现有最优模型。本工作将推动对自监督视频表征学习的深入理解。

原文摘要 · Abstract (English)

Self-supervised learning has emerged as a powerful paradigm for label-free model pretraining, particularly in the video domain, where manual annotation is costly and time-intensive. However, existing self-supervised approaches employ diverse experimental setups, making direct comparisons challenging due to the absence of a standardized benchmark. In this work, we establish a unified benchmark that enables fair comparisons across different methods. Additionally, we systematically investigate five critical aspects of self-supervised learning in videos: (1) dataset size, (2) model complexity, (3) data distribution, (4) data noise, and (5) feature representations. To facilitate this study, we evaluate six self-supervised learning methods across six network architectures, conducting extensive experiments on five benchmark datasets and assessing performance on two distinct downstream tasks. Our analysis reveals key insights into the interplay between pretraining strategies, dataset characteristics, pretext tasks, and model architectures. Furthermore, we extend these findings to Video Foundation Models (ViFMs), demonstrating their relevance in large-scale video representation learning. Finally, leveraging these insights, we propose a novel approach that significantly reduces training data requirements while surpassing state-of-the-art methods that rely on 10% more pretraining data. We believe this work will guide future research toward a deeper understanding of self-supervised video representation learning and its broader implications.

自监督视频表征基准评测模型优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。