提出新基准诊断多模态模型推理忠实地,发现两大失效模式。
SPD-Faith Bench: Diagnosing and Improving Faithfulness in Chain-of-Thought for Multimodal Large Language Models
- 基于细粒度图像差异推理构建诊断基准,强制显式视觉比对。
- 发现感知盲区与感知-推理脱节两大系统性失效,源于视觉注意力衰减。
- 提出无需训练的SAGE框架,提升视觉路由与推理一致性,适合研究可信AI者。
链式思维推理广泛用于提升多模态大语言模型(MLLMs)的可解释性,但其生成推理过程的忠实地仍不明确。以往研究主要关注感知幻觉,忽视了推理层面的不忠实问题。为将忠实地与语言先验分离,我们提出了SPD-Faith Bench,一个基于细粒度图像差异推理的诊断基准,强制进行显式视觉比较。对前沿MLLMs的评估揭示了两种系统性失效模式:感知盲区和感知-推理脱节。我们追溯这些失败源于残差流中的视觉注意力衰减与表征偏移。基于此分析,我们提出SAGE——一种无需训练的视觉证据校准框架,通过改善视觉路由并使推理与感知对齐。结果强调了超越回答正确性来显式评估忠实地的重要性。我们的基准与代码已公开于https://github.com/Johanson-colab/SPD-Faith-Bench。
原文摘要 · Abstract (English)
Chain-of-Thought reasoning is widely used to improve the interpretability of multimodal large language models (MLLMs), yet the faithfulness of the generated reasoning traces remains unclear. Prior work has mainly focused on perceptual hallucinations, leaving reasoning level unfaithfulness underexplored. To isolate faithfulness from linguistic priors, we introduce SPD-Faith Bench, a diagnostic benchmark based on fine-grained image difference reasoning that enforces explicit visual comparison. Evaluations on state-of-the-art MLLMs reveal two systematic failure modes, perceptual blindness and perception-reasoning dissociation. We trace these failures to decaying visual attention and representation shifts in the residual stream. Guided by this analysis, we propose SAGE, a train-free visual evidence-calibrated framework that improves visual routing and aligns reasoning with perception. Our results highlight the importance of explicitly evaluating faithfulness beyond response correctness. Our benchmark and codes are available at https://github.com/Johanson-colab/SPD-Faith-Bench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。