arXiv:2510.11196cs.CLcs.CV2025-10中稿 · ML4H 2025 Proceedi…被引 10

用多模态扰动评估医学视觉语言模型推理可靠性,发现答案准确不等于解释可信。

Evaluating Reasoning Faithfulness in Medical Vision-Language Models using Multimodal Perturbations

  • 通过文本与图像的临床、因果、信心三轴扰动,检验模型推理是否真实反映决策过程。
  • 6个模型中,开源模型解释质量普遍低于闭源,且文字线索比图像更易误导模型。
  • 仅靠最终答案准确率无法判断模型可信度,需关注解释与真实依据的一致性。

视觉语言模型(VLMs)常生成看似合理但与实际决策过程不符的链式思考(CoT)解释,影响其在高风险临床场景中的可信度。现有评估多关注答案准确性或格式规范,忽视这种错位。本文提出一个基于临床场景的胸部X光片视觉问答(VQA)框架,通过在临床真实性、因果归因和置信度校准三个维度上施加可控的文本与图像扰动,检测模型推理的忠实性。在由4名放射科医生参与的读者研究中,评估者与放射科医生之间的相关性在各维度均处于放射科医生间自然差异范围内:归因一致性较强(Kendall's $τ_b=0.670$),临床真实性中等($τ_b=0.387$),置信度语气较弱($τ_b=0.091$),需谨慎解读。对6个VLM的基准测试显示,答案准确率与解释质量可分离;即使模型察觉到人为注入线索,也不代表其解释是接地的;文本线索比视觉线索更显著改变解释内容。尽管部分开源模型在最终答案准确率上与闭源模型相当,但闭源模型在归因(25.0% vs. 1.4%)和临床真实性(36.1% vs. 31.7%)方面表现更优,凸显了仅依赖答案准确率评估的局限性及部署风险。

原文摘要 · Abstract (English)

Vision-language models (VLMs) often produce chain-of-thought (CoT) explanations that sound plausible yet fail to reflect the underlying decision process, undermining trust in high-stakes clinical use. Existing evaluations rarely catch this misalignment, prioritizing answer accuracy or adherence to formats. We present a clinically grounded framework for chest X-ray visual question answering (VQA) that probes CoT faithfulness via controlled text and image modifications across three axes: clinical fidelity, causal attribution, and confidence calibration. In a reader study (n=4), evaluator-radiologist correlations fall within the observed inter-radiologist range for all axes, with strong alignment for attribution (Kendall's $τ_b=0.670$), moderate alignment for fidelity ($τ_b=0.387$), and weak alignment for confidence tone ($τ_b=0.091$), which we report with caution. Benchmarking six VLMs shows that answer accuracy and explanation quality can be decoupled, acknowledging injected cues does not ensure grounding, and text cues shift explanations more than visual cues. While some open-source models match final answer accuracy, proprietary models score higher on attribution (25.0% vs. 1.4%) and often on fidelity (36.1% vs. 31.7%), highlighting deployment risks and the need to evaluate beyond final answer accuracy.

医学AI模型可信度视觉语言模型解释性评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。