区分医学模型推理与迎合行为,发现位置比来源更重要
Position, Not Provenance: Separating Reasoning Mediation from Sycophancy in Medical Vision-Language Models
- 通过修改模型自动生成的推理内容,测试其对预测的影响
- 强制续写比重新提示更可靠,位置决定推理是否被采纳
- 适合关注医疗AI可解释性与决策可信度的研究者
医学视觉语言模型(VLMs)在回答临床问题前会生成链式思维(CoT)推理,但这种推理是否真正影响预测尚不明确。我们提出CoT-Mediate行为框架,通过扰动模型自身生成推理中的单一临床属性,并测量预测是否随之改变。该框架结合双臂协议:比较重新提示证据与前缀强制续写,并采用溯源可控干预,仅改变相同推理的来源以分离推理影响与迎合行为。我们在LLaVA-Med和MedGemma上各评估1,000个VQA-RAD样本。结果表明,前缀强制续写始终比重新提示产生更高推理忠实度;溯源分析揭示模型存在特定依赖行为。移除视觉证据后,模型更依赖注入的推理,而侧位(laterality)是追踪最不准确的临床属性。结果表明,推理注入方式显著影响测量的忠实度,且上下文位置而非声明来源,是决定模型是否采纳推理的关键因素。
原文摘要 · Abstract (English)
Medical vision-language models (VLMs) generate chain-of-thought (CoT) reasoning before answering clinical questions, but whether this reasoning causally influences predictions remains unclear. We present CoT-Mediate, a behavioral framework that perturbs a single clinically meaningful attribute within a model's own generated reasoning and measures whether the resulting prediction follows the edited reasoning. Our framework combines a dual-arm protocol comparing re-prompted evidence with prefix-forced continuation, together with a provenance-controlled intervention that varies only the attributed source of identical reasoning to disentangle reasoning mediation from sycophancy. We evaluate LLaVA-Med and MedGemma on 1,000 VQA-RAD samples each. Prefix-forced continuation consistently yields higher mediation faithfulness than re-prompting, while the provenance analysis reveals distinct model-specific deference behaviors. Across both models, removing visual evidence increases reliance on injected reasoning, whereas laterality is the least faithfully tracked clinical attribute. These results show that the mechanism used to inject reasoning substantially affects measured faithfulness and that contextual position, rather than stated provenance, is the primary determinant of whether medical VLMs use their generated reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。