arXiv:2512.12218cs.CVcs.CL2025-12Conference of the …被引 4

提出视觉忠实度评估新维度,让模型推理更靠谱

Journey Before Destination: On the importance of Visual Faithfulness in Slow Thinking

  • 分离推理链中的感知与推理步骤,用现成模型判断感知是否真实
  • 无需训练即可检测并修复不忠实的视觉感知,准确率不变
  • 适合关注多模态模型可信度的研究者和开发者

推理增强型视觉语言模型(VLMs)通过显式推理链提升能力与透明性,但也出现新问题:模型可能通过不真实的中间步骤得出正确答案,或虽推理忠实却最终失败。现有仅评估最终答案准确性的方法无法区分这些行为。本文提出将推理链的视觉忠实度作为独立评估维度,关注其感知步骤是否基于图像。我们设计了一种无需训练和参考的框架,将推理链分解为感知与推理步骤,并利用现成的VLM评判器进行逐步评估,同时通过人工元评估验证有效性。在此基础上,提出轻量级自反思机制,可检测并局部重生成不忠实的感知步骤,无需训练。在多个推理训练的VLM及感知密集型基准上,该方法显著降低非忠实感知率,同时保持最终答案准确率,提升了多模态推理的可靠性。

原文摘要 · Abstract (English)

Reasoning-augmented vision language models (VLMs) generate explicit chains of thought that promise greater capability and transparency but also introduce new failure modes: models may reach correct answers via visually unfaithful intermediate steps, or reason faithfully yet fail on the final prediction. Standard evaluations that only measure final-answer accuracy cannot distinguish these behaviors. We introduce the visual faithfulness of reasoning chains as a distinct evaluation dimension, focusing on whether the perception steps of a reasoning chain are grounded in the image. We propose a training- and reference-free framework that decomposes chains into perception versus reasoning steps and uses off-the-shelf VLM judges for step-level faithfulness, additionally verifying this approach through a human meta-evaluation. Building on this metric, we present a lightweight self-reflection procedure that detects and locally regenerates unfaithful perception steps without any training. Across multiple reasoning-trained VLMs and perception-heavy benchmarks, our method reduces Unfaithful Perception Rate while preserving final-answer accuracy, improving the reliability of multimodal reasoning.

多模态推理视觉忠实度自反思

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。