arXiv:2411.17474cs.CV2024-11CVPR被引 7

首次系统评估自监督模型的中层视觉能力,发现其与高层任务表现关联弱。

Probing the Mid-level Vision Capabilities of Self-Supervised Learning

  • 构建8项中层视觉任务基准,评估22个SSL模型
  • 发现中层与高层任务性能相关性弱,部分模型表现失衡
  • 揭示预训练目标与网络结构对中层能力的关键影响

中层视觉能力(如通用物体定位、三维几何理解)不仅在人类视觉发展中早期出现,也是计算机视觉诸多实际应用的核心。然而,当前自监督学习(SSL)方法主要针对高层识别任务设计和评估,对其在中层视觉方面的能力关注不足。本文引入一套基准评测协议,对22种主流SSL模型在8项中层视觉任务上进行系统性、受控评估。实验表明,中层任务与高层任务表现之间存在弱相关性;部分方法在两类任务间表现严重失衡,也有少数模型在两者中均表现优异。我们进一步分析了预训练目标与网络架构等关键因素对中层视觉能力的影响。本研究为理解SSL模型所学内容提供了全面视角,补充了以高层任务为主的现有研究。希望未来研究能将中层视觉能力纳入评估体系。

原文摘要 · Abstract (English)

Mid-level vision capabilities - such as generic object localization and 3D geometric understanding - are not only fundamental to human vision but are also crucial for many real-world applications of computer vision. These abilities emerge with minimal supervision during the early stages of human visual development. Despite their significance, current self-supervised learning (SSL) approaches are primarily designed and evaluated for high-level recognition tasks, leaving their mid-level vision capabilities largely unexamined. In this study, we introduce a suite of benchmark protocols to systematically assess mid-level vision capabilities and present a comprehensive, controlled evaluation of 22 prominent SSL models across 8 mid-level vision tasks. Our experiments reveal a weak correlation between mid-level and high-level task performance. We also identify several SSL methods with highly imbalanced performance across mid-level and high-level capabilities, as well as some that excel in both. Additionally, we investigate key factors contributing to mid-level vision performance, such as pretraining objectives and network architectures. Our study provides a holistic and timely view of what SSL models have learned, complementing existing research that primarily focuses on high-level vision tasks. We hope our findings guide future SSL research to benchmark models not only on high-level vision tasks but on mid-level as well.

自监督学习中层视觉模型评估视觉理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。