arXiv:2607.28211cs.CV2026-07

模型越大越不抗偏见,数据质量比规模更重要。

Scaling Vision-Language Models Is Not Enough to Mitigate Bias

  • 对比194个VLM模型,发现规模与抗偏见能力相关性随评估复杂度下降。
  • 在多属性偏见任务中,模型规模与性能相关性近乎为零(ρ=0.05)。
  • 精心筛选的数据集可提升最差群体准确率25%,适合关注公平性的研究者。

视觉语言模型(如CLIP)已成为多模态系统的基础,但其对虚假关联的鲁棒性在大规模下仍不明确。我们开展了首个大规模实证研究,涵盖194个公开VLMs,包括16个模型家族,覆盖多种模型规模、24个训练数据集及三个评估基准:ImageNet(整体性能)、CelebA(典型单属性偏见)和UrbanCars(复杂多属性偏见)。结果显示,模型规模与性能的斯皮尔曼相关系数从ImageNet的ρ=0.68下降至单属性偏见的ρ=0.48,进一步降至多属性偏见的ρ=0.05。相比之下,训练数据的规模与质量在两类偏见基准中与最差群体准确率呈现更稳定的关系。值得注意的是,经过人工筛选的数据集在相同规模下可使最差群体准确率提升最高达25%。此外,架构选择(如图像块大小、分辨率)的影响高度依赖于具体场景,取决于偏见类型及其在图像中的空间分布。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) such as CLIP are now foundational to multimodal systems, yet their robustness to spurious correlations remains poorly understood at scale. We present the first large-scale empirical study of 194 publicly available VLMs, including 16 model families, covering a wide range of model sizes, 24 training datasets, and three evaluation benchmarks, namely ImageNet (overall performance), CelebA (typical single-attribute bias), and UrbanCars (complex multi-attribute biases). Across these settings, the Spearman correlation between model scale and performance weakens as evaluation shifts from ImageNet ($ρ{=}0.68$) to single-attribute ($ρ{=}0.48$) and further to multi-attribute ($ρ{=}0.05$) bias benchmarks. In contrast, properties of the training data (size and quality) show more consistent relationships with worst-group accuracy across both bias benchmarks. Notably, curated datasets yield improvements of up to 25% over uncurated alternatives at a comparable scale. Finally, the effect of architectural choices (e.g., patch size, image resolution) is highly context-dependent, varying with the nature of the benchmark, including the type of bias and its spatial distribution within images.

偏见缓解多模态数据质量模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。