arXiv:2507.07985cs.CV2025-07中稿 · GCPR 2025被引 3

发现自然数据属性导致CLIP无法正确绑定物体与属性

Common Data Properties Limit Object-Attribute Binding in CLIP

  • 用合成数据实验揭示自然数据的低属性密度等特性影响绑定能力
  • 仅当数据具备特定属性时,CLIP才能实现接近完美的物体-属性绑定
  • 挑战了加大批次或人工构造难例可改善绑定的普遍认知

对比视觉语言模型如CLIP被广泛用于零样本分类或多模态模型的视觉编码器,但其表征存在明显局限。例如,CLIP学习的是词袋表示,难以区分“一艘黄色潜艇和一辆蓝色公交车”与“一辆蓝色潜艇和一艘黄色公交车”。以往尝试通过训练中加入难例或修改架构来解决此问题,但未能根本解决。本文认为,解决CLIP绑定问题的关键在于数据本身。通过构建合成数据集,系统研究数据属性对绑定能力的影响,发现自然数据中的低属性密度、不完整描述及显著性偏差(即标注者更倾向描述最显著对象)会严重损害绑定性能。相反,尽管增加批量大小或显式构造难例,也无法使CLIP实现可靠绑定。只有在数据充分表达这些关键属性时,CLIP才能实现近乎完美的绑定。

原文摘要 · Abstract (English)

Contrastive vision-language models like CLIP are used for a large variety of applications, such as zero-shot classification or as vision encoder for multi-modal models. Despite their popularity, their representations show major limitations. For instance, CLIP models learn bag-of-words representations and, as a consequence, fail to distinguish whether an image is of ``a yellow submarine and a blue bus'' or ``a blue submarine and a yellow bus''. Previous attempts to fix this issue added hard negatives during training or modified the architecture, but failed to resolve the problem in its entirety. We suspect that the missing insights to solve the binding problem for CLIP are hidden in arguably the most important part of learning algorithms: the data. In this work, we fill this gap by rigorously identifying the influence of data properties on CLIP's ability to learn binding using a synthetic dataset. We find that common properties of natural data such as low attribute density, incomplete captions, and the saliency bias, a tendency of human captioners to describe the object that is ``most salient'' to them, have a detrimental effect on binding performance. In contrast to common belief, we find that neither scaling the batch size, i.e., implicitly adding more hard negatives, nor explicitly creating hard negatives enables CLIP to learn reliable binding. Only when the data expresses our identified data properties does CLIP learn almost perfect binding.

CLIP视觉语言模型绑定问题数据属性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。