arXiv:2509.12265cs.CVcs.AI2025-09

研究大模型在图像分类中的简单性偏好,发现其与任务特性密切相关。

A Modern Look at Simplicity Bias in Image Classification Tasks

  • 提出新的频率感知度量方法,更精细捕捉大模型的简单性偏好差异。
  • 实验证明强简单性偏好利于分布外泛化,但不利于对抗鲁棒性。
  • 适合关注模型归纳偏置与任务匹配的研究者阅读。

神经网络的简单性偏好(SB)是其泛化能力的关键因素。近期研究表明,过度的简单性偏好可能损害复杂任务的表现,且该偏好需求随任务变化。现有研究多聚焦于简单模型或合成任务,对大型模型中简单性偏好的测量仍具挑战,且其与各类图像分类任务的相关性尚不明确。本文探究了CLIP模型的简单性偏好与其在多种图像分类任务上的性能关系。首先,理论分析了现有复杂度度量在小型模型中的局限性,并提出一种频率感知的度量方法,能更精细地捕捉简单性偏好的差异。在两种近期的简单性调节方法下验证该度量,结果表明其比传统方法更具信息量和一致性。其次,考察模型简单性偏好与多项图像分类任务性能的关系,涵盖零样本和微调设置。实验揭示多样行为模式:更强的简单性偏好有助于分布外泛化,但不利于对抗鲁棒性。这些结果凸显了将模型归纳偏置与目标任务特征对齐的重要性。

原文摘要 · Abstract (English)

The simplicity Bias (SB) of neural networks, i.e.\ their tendency to represent simple functions, is a key factor in their generalization capabilities. Recent studies show that an excessive SB may harm performance on complex tasks, and the need for this bias varies across tasks. Many of these studies focus on simple models or synthetic tasks. It remains challenging to measure the SB in large models and little is known about the relevance of the SB to various image classification tasks. In this paper, we investigate the relationship between the SB in CLIP models and their performance across image classification tasks. First, we theoretically analyze the potential limitation of existing measures of complexity that have been used to characterize small models. To address this, we propose a frequency-aware measure capturing finer-grained SB differences. We validate this measure on CLIP models subjected to two recent SB-modulation methods, demonstrating that it is more informative and consistent than previous measures. Second, we examine the relation between the SB of those models and their performance across a range of image classification tasks, including zero-shot and fine-tuning settings. These experiments reveal a range of behaviors. For example, a stronger SB correlates with a better performance on OOD generalization than on adversarial robustness. These results highlight the benefits of aligning a model's inductive biases with the characteristics of the target task.

简单性偏好CLIP图像分类归纳偏置

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。