arXiv:2506.16950cs.CVcs.LG2025-06ICML被引 6

为网页规模视觉模型设计新基准,挑战现有模型的分布外泛化能力。

LAION-C: An Out-of-Distribution Benchmark for Web-Scale Vision Models

  • 提出六种全新分布外扭曲类型,确保对网页数据集仍具挑战性。
  • 顶尖模型在新基准上表现显著下降,体现真实泛化能力瓶颈。
  • 首次实现人与模型在分布外任务上的直接对比,揭示模型超越人类趋势。

分布外(OOD)鲁棒性是计算机视觉模型的重要特性。提升模型鲁棒性需依赖高质量的鲁棒性基准来量化进展。尽管图像网(ImageNet)时代提出了多种基准数据集如ImageNet-C,但其中多数畸变类型已不再属于分布外——因当前大规模网络抓取数据集(如LAION)本身已包含模糊、JPEG压缩等常见失真。因此,这些基准已不适配用于评估现代网页规模数据集下的分布外鲁棒性。事实上,近期模型在旧基准上得分趋于饱和,表明无法判断模型在网页数据训练后是否真正提升了分布外泛化能力,还是仅在训练中接触过测试失真。为此,我们引入LAION-C作为ImageNet-C的替代基准。LAION-C包含六种专为分布外设计的新畸变类型,即使在大型网页数据集如LAION中也保持分布外特性。对前沿模型的全面评估显示,该数据集对当代模型构成严峻挑战,包括多模态大模型(如Gemini和GPT-4o)。我们还进行了心理物理学实验,评估人类观察者对这些畸变的感知难度,从而实现模型与实验室级人类鲁棒性数据的对比。结果发现分布外泛化出现范式转变:从人类优于模型,转向最优模型已可匹配甚至超越最优人类观察者。

原文摘要 · Abstract (English)

Out-of-distribution (OOD) robustness is a desired property of computer vision models. Improving model robustness requires high-quality signals from robustness benchmarks to quantify progress. While various benchmark datasets such as ImageNet-C were proposed in the ImageNet era, most ImageNet-C corruption types are no longer OOD relative to today's large, web-scraped datasets, which already contain common corruptions such as blur or JPEG compression artifacts. Consequently, these benchmarks are no longer well-suited for evaluating OOD robustness in the era of web-scale datasets. Indeed, recent models show saturating scores on ImageNet-era OOD benchmarks, indicating that it is unclear whether models trained on web-scale datasets truly become better at OOD generalization or whether they have simply been exposed to the test distortions during training. To address this, we introduce LAION-C as a benchmark alternative for ImageNet-C. LAION-C consists of six novel distortion types specifically designed to be OOD, even for web-scale datasets such as LAION. In a comprehensive evaluation of state-of-the-art models, we find that the LAION-C dataset poses significant challenges to contemporary models, including MLLMs such as Gemini and GPT-4o. We additionally conducted a psychophysical experiment to evaluate the difficulty of our corruptions for human observers, enabling a comparison of models to lab-quality human robustness data. We observe a paradigm shift in OOD generalization: from humans outperforming models, to the best models now matching or outperforming the best human observers.

分布外视觉模型基准测试鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。