通过新数据集和分组反义学习,提升视觉语言模型对否定语义的识别能力。
Mastering Negation: Boosting Grounding Models via Grouped Opposition-Based Learning
- 构建包含正反语义标注的新数据集D-Negation,支持否定推理训练。
- 在少量参数微调下,正负语义识别分别提升4.4和5.7 mAP。
- 适合关注视觉定位、自然语言理解的科研与工程人员。
当前视觉-语言检测与定位模型多聚焦于正面语义提示,难以准确解析含否定语义的复杂表达。主要原因在于缺乏高质量的训练数据,无法充分捕捉具有区分性的否定样本及否定感知的语言描述。为此,我们提出D-Negation数据集,为物体提供正负语义双重标注。基于否定推理在自然语言中频繁出现的现象,我们进一步设计了分组反义学习框架,从有限样本中学习否定感知表征。该方法将D-Negation中的对立语义描述组织成结构化组,并设计两种互补损失函数,促使模型理解否定与语义限定词。我们将该数据集与学习策略集成至先进语言引导定位模型中。仅微调少于10%的模型参数,正负语义评估分别取得最高4.4 mAP和5.7 mAP的提升。结果表明,显式建模否定语义可显著增强视觉-语言定位模型的鲁棒性与定位精度。
原文摘要 · Abstract (English)
Current vision-language detection and grounding models predominantly focus on prompts with positive semantics and often struggle to accurately interpret and ground complex expressions containing negative semantics. A key reason for this limitation is the lack of high-quality training data that explicitly captures discriminative negative samples and negation-aware language descriptions. To address this challenge, we introduce D-Negation, a new dataset that provides objects annotated with both positive and negative semantic descriptions. Building upon the observation that negation reasoning frequently appears in natural language, we further propose a grouped opposition-based learning framework that learns negation-aware representations from limited samples. Specifically, our method organizes opposing semantic descriptions from D-Negation into structured groups and formulates two complementary loss functions that encourage the model to reason about negation and semantic qualifiers. We integrate the proposed dataset and learning strategy into a state-of-the-art language-based grounding model. By fine-tuning fewer than 10 percent of the model parameters, our approach achieves improvements of up to 4.4 mAP and 5.7 mAP on positive and negative semantic evaluations, respectively. These results demonstrate that explicitly modeling negation semantics can substantially enhance the robustness and localization accuracy of vision-language grounding models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。