arXiv:2501.15431cs.CVcs.AI2025-01中稿 · publication in the…被引 3

检验自监督模型在ImageNet上的微小提升能否迁移到相似数据集

Self-supervised Benchmark Lottery on ImageNet: Do Marginal Improvements Translate to Improvements on Similar Datasets?

  • 测试12种自监督框架在5个ImageNet变体上的表现
  • DINO和Swav在相似数据集上性能大幅下降,MoCo和Barlow Twins更稳定
  • 建议用多数据集统一指标避免单一数据集的'基准彩票'陷阱

机器学习研究依赖基准来评估新模型的有效性。近期有研究指出,许多仅带来小幅性能提升的模型,实际上是靠‘基准彩票’获胜。图像识别领域的重要基准ImageNet常被用于展示新模型性能。近年来,大量自监督学习(SSL)框架在ImageNet上取得微小进步,但这些改进是否能在类似数据集上延续?本文评估了12种主流SSL框架在5个ImageNet变体上的表现,发现原本在ImageNet上表现优异的模型(如DINO、Swav)在相似数据集上性能显著下降,而MoCo与Barlow Twins则表现出更强的稳定性。这表明,仅在ImageNet验证集上评估模型会掩盖其真实能力,因此我们呼吁采用包含多个ImageNet变体的统一评估指标,以实现更公平的基准测试。

原文摘要 · Abstract (English)

Machine learning (ML) research strongly relies on benchmarks in order to determine the relative effectiveness of newly proposed models. Recently, a number of prominent research effort argued that a number of models that improve the state-of-the-art by a small margin tend to do so by winning what they call a "benchmark lottery". An important benchmark in the field of machine learning and computer vision is the ImageNet where newly proposed models are often showcased based on their performance on this dataset. Given the large number of self-supervised learning (SSL) frameworks that has been proposed in the past couple of years each coming with marginal improvements on the ImageNet dataset, in this work, we evaluate whether those marginal improvements on ImageNet translate to improvements on similar datasets or not. To do so, we investigate twelve popular SSL frameworks on five ImageNet variants and discover that models that seem to perform well on ImageNet may experience significant performance declines on similar datasets. Specifically, state-of-the-art frameworks such as DINO and Swav, which are praised for their performance, exhibit substantial drops in performance while MoCo and Barlow Twins displays comparatively good results. As a result, we argue that otherwise good and desirable properties of models remain hidden when benchmarking is only performed on the ImageNet validation set, making us call for more adequate benchmarking. To avoid the "benchmark lottery" on ImageNet and to ensure a fair benchmarking process, we investigate the usage of a unified metric that takes into account the performance of models on other ImageNet variant datasets.

自监督学习基准测试迁移性能ImageNet

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。