arXiv:2604.13313cs.LG2026-04被引 1

通过强化具体词汇的对比负样本提升视觉语言模型组合理解能力

Concrete Jungle: Towards Concreteness Paved Contrastive Negative Mining for Compositional Understanding

  • 以词汇具体性指导负样本生成,增强语义差异
  • 提出Cement损失缓解梯度失衡,提升细微语义学习
  • 适合关注视觉语言模型推理能力提升的研究者

视觉-语言模型在组合推理上表现脆弱,尤其对词序和属性绑定敏感。这源于对比预训练中缺乏能区分细微语义差异的有信息量样本。尽管硬负样本挖掘具潜力,但现有方法未明确指定应修改哪些语言成分。本文将词汇具体性确立为负样本有效性的根本决定因素:修改高具体性词语可产生更显著的结构与视觉差异,提供更强学习信号。据此提出ConcretePlant,系统性地识别并操纵感知基础概念。对InfoNCE的分析揭示严重梯度失衡,易区分样本过度主导优化过程,限制了细微学习的容量。为此提出基于边际的Cement损失,通过关联心理语言学评分与样本难度,动态调节每对训练样本的惩罚力度。全面评估证实理论假设。集成框架Slipform在多个组合评估基准、跨模态检索及单/多标签线性探测任务中均达到当前最优性能。

原文摘要 · Abstract (English)

Vision-Language Models demonstrate remarkable capabilities but often struggle with compositional reasoning, exhibiting vulnerabilities regarding word order and attribute binding. This limitation arises from a scarcity of informative samples needed to differentiate subtle semantic variations during contrastive pretraining. Although hard negative mining offers a promising remedy, existing methods lack explicit mechanisms to dictate which linguistic elements undergo modification. Instead of engineering generative architectures, this study establishes lexical concreteness as a fundamental determinant of negative sample efficacy. Modifying highly concrete terms generates more pronounced structural and visual discrepancies, providing a substantially stronger learning signal. Leveraging this principle, ConcretePlant is proposed to systematically isolate and manipulate perceptually grounded concepts. Analyses of the InfoNCE further reveals a severe gradient imbalance, where easily distinguishable pairs disproportionately overwhelm the optimization process and restrict the bandwidth available for nuanced learning. To resolve this degradation, the Cement loss is formulated utilizing a margin-based approach. By correlating psycholinguistic scores with sample difficulty, this objective dynamically calibrates the penalization applied to individual training pairs. Comprehensive evaluations substantiate these theoretical claims. The integrated framework, designated as Slipform, achieves state-of-the-art accuracy across diverse compositional evaluation benchmarks, general cross-modal retrieval, single and multi label linear probing.

视觉语言对比学习组合推理负样本挖掘

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。