arXiv:2506.08227cs.CV2025-06被引 4

发现主流视觉语言基准存在固有偏差,可能误导模型评估

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks

  • 分析17个常用基准的设计缺陷,揭示数据构造中的系统性偏差
  • 简单启发式方法表现媲美CLIP模型,说明基准无法真实衡量组合理解能力
  • 提出构建更鲁棒基准的建议,适合评估与设计视觉语言模型的研究者

我们研究了17个常用于评估视觉语言模型(VLMs)组合理解能力的基准(如SugarCREPE、VALSE),审视其数据来源(如MS-COCO)和筛选流程(如构造负样本图像/描述),发现大多数基准存在内在偏差。研究发现,仅依赖词长或语言模型对数似然等盲启发式方法的表现可与CLIP模型相当,表明这些基准未能有效测量真正的组合理解能力。我们证明根本原因在于正负样本间分布不对称,由基准构建过程导致。为此,我们提出若干关键建议,以构建更稳健、不易被简单策略攻破的视觉语言组合理解基准。

原文摘要 · Abstract (English)

We investigate 17 benchmarks (e.g. SugarCREPE, VALSE) commonly used for measuring compositional understanding capabilities of vision-language models (VLMs). We scrutinize design choices in their construction, including data source (e.g. MS-COCO) and curation procedures (e.g. constructing negative images/captions), uncovering several inherent biases across most benchmarks. We find that blind heuristics (e.g. token-length, log-likelihood under a language model) perform on par with CLIP models, indicating that these benchmarks do not effectively measure compositional understanding. We demonstrate that the underlying factor is a distribution asymmetry between positive and negative images/captions, induced by the benchmark construction procedures. To mitigate these issues, we provide a few key recommendations for constructing more robust vision-language compositional understanding benchmarks, that would be less prone to such simple attacks.

视觉语言基准评估组合理解偏差分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。