arXiv:2505.23945cs.CLcs.AI2025-05EMNLP被引 16

首次系统研究视觉语言模型推理忠实度,发现图像偏见常被忽略,且存在突然改答案的不一致现象。

A Closer Look at Bias and Chain-of-Thought Faithfulness of Large (Vision) Language Models

  • 构建细粒度评估框架,区分文本与图像偏见对推理的影响
  • 图像隐含偏见极少被模型识别和说明,远低于文本显性偏见
  • 发现模型存在'不一致推理':先正确后突变答案,或为偏见预警信号

链式思维(CoT)推理虽能提升大语言模型性能,但其推理过程是否真实反映模型内部逻辑仍存疑问。本文首次全面研究大型视觉语言模型(LVLMs)中CoT的忠实度,分析文本和此前未被关注的图像偏见如何影响推理及偏见表达。我们提出一种新颖的细粒度评估流程,可更精准分类偏见表达模式,揭示模型在处理不同类型偏见时的关键差异。结果显示,即使在专门设计用于推理的模型中,细微图像偏见也极少被明确指出,远低于明显文本偏见。此外,许多模型表现出一种此前未被识别的现象——‘不一致’推理:先正确推理,随后突然改变答案,或可作为检测非忠实推理的早期警示。该评估框架亦用于重审语言模型在不同隐含线索下的推理忠实度,发现当前纯语言模型仍难以揭示未明确陈述的线索。

原文摘要 · Abstract (English)

Chain-of-thought (CoT) reasoning enhances performance of large language models, but questions remain about whether these reasoning traces faithfully reflect the internal processes of the model. We present the first comprehensive study of CoT faithfulness in large vision-language models (LVLMs), investigating how both text-based and previously unexplored image-based biases affect reasoning and bias articulation. Our work introduces a novel, fine-grained evaluation pipeline for categorizing bias articulation patterns, enabling significantly more precise analysis of CoT reasoning than previous methods. This framework reveals critical distinctions in how models process and respond to different types of biases, providing new insights into LVLM CoT faithfulness. Our findings reveal that subtle image-based biases are rarely articulated compared to explicit text-based ones, even in models specialized for reasoning. Additionally, many models exhibit a previously unidentified phenomenon we term ``inconsistent'' reasoning - correctly reasoning before abruptly changing answers, serving as a potential canary for detecting biased reasoning from unfaithful CoTs. We then apply the same evaluation pipeline to revisit CoT faithfulness in LLMs across various levels of implicit cues. Our findings reveal that current language-only reasoning models continue to struggle with articulating cues that are not overtly stated.

视觉语言模型链式思维偏见检测推理忠实度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。