arXiv:2504.08974cs.AIcs.CV2025-04EMNLP被引 7

研究视觉语言模型在图文冲突时的推理偏见及应对策略

Mixed Signals: Decoding VLMs' Reasoning and Underlying Bias in Vision-Language Conflict

  • 通过构建五组图文不一致数据集,分析模型在冲突场景下的决策倾向
  • 模型在简单任务中更信文本,复杂任务中转向图像,偏见幅度达±74.4%
  • 提示工程与分步分析可缓解偏见,效果依赖任务难度和模型能力

视觉语言模型(VLMs)在整合视觉与文本信息方面表现优异,但其跨模态推理机制尚不明确。本文通过分析模型在图文冲突场景下的行为,揭示其内在偏见。为此,我们扩展现有基准,构建了涵盖数学、科学和视觉描述的五个图文不匹配数据集。分析显示,模型在简单查询中更倾向文本,复杂任务中则偏向图像,二者偏好差异达+56.8%(图像偏好)至-74.4%(文本偏好)。此外,我们测试了三种缓解策略:简单提示修改、显式引导冲突处理(类链式思考提示),以及先分别分析再融合结果的任务分解法。结果显示,这些策略的有效性因任务复杂度、模型性能及具体模态而异。

原文摘要 · Abstract (English)

Vision-language models (VLMs) have demonstrated impressive performance by effectively integrating visual and textual information to solve complex tasks. However, it is not clear how these models reason over the visual and textual data together, nor how the flow of information between modalities is structured. In this paper, we examine how VLMs reason by analyzing their biases when confronted with scenarios that present conflicting image and text cues, a common occurrence in real-world applications. To uncover the extent and nature of these biases, we build upon existing benchmarks to create five datasets containing mismatched image-text pairs, covering topics in mathematics, science, and visual descriptions. Our analysis shows that VLMs favor text in simpler queries but shift toward images as query complexity increases. This bias correlates with model scale, with the difference between the percentage of image- and text-preferred responses ranging from +56.8% (image favored) to -74.4% (text favored), depending on the task and model. In addition, we explore three mitigation strategies: simple prompt modifications, modifications that explicitly instruct models on how to handle conflicting information (akin to chain-of-thought prompting), and a task decomposition strategy that analyzes each modality separately before combining their results. Our findings indicate that the effectiveness of these strategies in identifying and mitigating bias varies significantly and is closely linked to the model's overall performance on the task and the specific modality in question.

视觉语言模型推理偏见多模态提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。