arXiv:2504.05706cs.CV2025-04被引 3

对比12种视频自监督模型,发现无一能全面适应真实场景变化。

SEVERE++: Evaluating Benchmark Sensitivity in Generalization of Video Representation Learning

  • 构建跨领域、小样本、细粒度和多任务的综合评估框架
  • 1100次实验显示:视觉-文本模型虽大但泛化能力不优
  • 新基准揭示模型在真实场景中的脆弱性,适合研究泛化能力者参考

自监督学习在视频表征学习中取得显著进展,提供了无需人工标注的可扩展替代方案。尽管在标准动作识别基准上表现良好,现有方法大多仅在Kinetics-400预训练并微调于相似数据集,限制了对真实世界泛化能力的理解。本文全面评估现代视频自监督模型,聚焦四个关键下游因素:领域偏移、样本效率、动作粒度和任务多样性。基于先前对CNN对比学习的基准敏感性分析,本研究扩展至覆盖最先进的基于Transformer的纯视频与视频-文本模型。具体包含12种基于Transformer的方法(7种纯视频,5种视频-文本),并与10种基于CNN的方法进行比较,总计在8个数据集和7个下游任务上完成超过1100次实验。结果表明,尽管架构进步,基于Transformer的模型仍对下游条件敏感;无一种方法能在所有因素下一致表现良好:纯视频Transformer在领域偏移下表现更优,CNN在细粒度任务中胜出,而视频-文本模型尽管大规模预训练却常表现不佳。此外,近期的Transformer模型并未持续优于早期方法。研究结果为当前视频自监督学习方法的优劣提供了详尽视图,并建立了一个统一的泛化评估基准。

原文摘要 · Abstract (English)

Continued advances in self-supervised learning have led to significant progress in video representation learning, offering a scalable alternative to supervised approaches by removing the need for manual annotations. Despite strong performance on standard action recognition benchmarks, video self-supervised learning methods are largely evaluated under narrow protocols, typically pretraining on Kinetics-400 and fine-tuning on similar datasets, limiting our understanding of their generalization in real world scenarios. In this work, we present a comprehensive evaluation of modern video self-supervised models, focusing on generalization across four key downstream factors: domain shift, sample efficiency, action granularity, and task diversity. Building on our prior work analyzing benchmark sensitivity in CNN-based contrastive learning, we extend the study to cover state-of-the-art transformer-based video-only and video-text models. Specifically, we benchmark 12 transformer-based methods (7 video-only, 5 video-text) and compare them to 10 CNN-based methods, totaling over 1100 experiments across 8 datasets and 7 downstream tasks. Our analysis shows that, despite architectural advances, transformer-based models remain sensitive to downstream conditions. No method generalizes consistently across all factors, video-only transformers perform better under domain shifts, CNNs outperform for fine-grained tasks, and video-text models often underperform despite large scale pretraining. We also find that recent transformer models do not consistently outperform earlier approaches. Our findings provide a detailed view of the strengths and limitations of current video SSL methods and offer a unified benchmark for evaluating generalization in video representation learning.

视频表征自监督学习泛化能力基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。