用自然语言推理增强视觉语言模型的组合推理能力
Natural Language Inference Improves Compositionality in Vision-Language Models
- 基于自然语言蕴含关系生成语义一致的扩展句,提升多样性
- 在Winoground和EqBen上分别提升19.2%和12.9%的准确率
- 无需额外微调,适合提升模型对图像文本对齐的理解
视觉语言模型在组合推理方面仍具挑战,常难以正确关联物体、属性与空间关系。现有方法依赖大语言模型对文本描述进行问答拆解,但多停留在表面语义,缺乏深层词汇理解,并引入错误假设。为此,我们提出基于矛盾与蕴含的标题扩展方法(CECE),利用自然语言推理从原始前提生成蕴含与矛盾句,在保持核心语义的同时增加词汇多样性。大量实验表明,CECE提升了模型可解释性,减少了对偏差或表层特征的依赖。通过平衡原始前提,该方法在无需额外微调的情况下显著优于先前方法,在衡量图像-文本对齐的人类判断一致性任务上达到新纪录,于Winoground上组得分提升19.2%,在EqBen上提升12.9%(相较于最佳前序工作,后者使用针对性数据微调)。
原文摘要 · Abstract (English)
Compositional reasoning in Vision-Language Models (VLMs) remains challenging as these models often struggle to relate objects, attributes, and spatial relationships. Recent methods aim to address these limitations by relying on the semantics of the textual description, using Large Language Models (LLMs) to break them down into subsets of questions and answers. However, these methods primarily operate on the surface level, failing to incorporate deeper lexical understanding while introducing incorrect assumptions generated by the LLM. In response to these issues, we present Caption Expansion with Contradictions and Entailments (CECE), a principled approach that leverages Natural Language Inference (NLI) to generate entailments and contradictions from a given premise. CECE produces lexically diverse sentences while maintaining their core meaning. Through extensive experiments, we show that CECE enhances interpretability and reduces overreliance on biased or superficial features. By balancing CECE along the original premise, we achieve significant improvements over previous methods without requiring additional fine-tuning, producing state-of-the-art results on benchmarks that score agreement with human judgments for image-text alignment, and achieving an increase in performance on Winoground of +19.2% (group score) and +12.9% on EqBen (group score) over the best prior work (finetuned with targeted data).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。