arXiv:2603.28387cs.AIcs.LG2026-03中稿 · EMNLP

小模型靠影像提示提升表现,但可能误信虚构诊断信息。

Prompts Without Evidence: How Neuroimaging Mentions Shift Clinical Vision-Language Model Predictions

  • 用影像上下文提示诱导小模型性能跃升
  • 加入影像描述后小模型F1最高提升0.66
  • 适合关注临床AI幻觉风险的研究者

可信的临床AI应基于真实证据,避免依赖表面特征。我们评估了12个开源视觉语言模型(VLMs)在两个临床神经影像队列上的二分类任务,分别用于情感障碍和认知衰退判断。两个队列均包含按原始研究协议采集的结构磁共振成像(MRI)。先前研究未证实这些影像可作为独立诊断依据。然而,引入影像上下文后,较小的VLMs在增强条件下F1最高提升0.66,达到与大一个数量级的模型相当的性能。置信度分析显示,多数校准改善发生在加入MRI参考提示后、图像输入前。初步专家案例研究发现,所有条件下模型忠实性仍较低,且引入未经验证的临床细节。单模型干预中,偏好对齐虽抑制了引用MRI的行为,但削弱了增强条件下的优势,问题未根本解决。结果警示:不可将表面指标提升视为真正多模态融合的证据,对临床VLM部署具有直接启示。

原文摘要 · Abstract (English)

Trustworthy clinical AI must use real evidence and avoid relying on surface-level artifacts. We evaluate 12 open-weight vision-language models (VLMs) on two clinical neuroimaging cohorts for binary classification of affective disorders and cognitive decline. Both cohorts include structural magnetic resonance imaging (MRI) acquired under their original research protocols. Prior work does not establish the included neuroimaging inputs as reliable stand-alone diagnostic evidence for the present tasks. Nevertheless, when neuroimaging context is introduced, smaller VLMs gain up to 0.66 F1 under the evaluated augmented conditions, becoming competitive with models an order of magnitude larger. Confidence estimation shows that most of the calibration improvement for the analyzed smaller models occurs after the MRI reference is added to the prompt, before any image is supplied. Our preliminary expert case study finds that faithfulness remains low in every condition examined, with the reviewed model introducing unverified clinical details. Finally, in our single-model intervention, preference alignment suppresses MRI-referencing behavior but reduces the augmented-condition advantage, leaving the underlying issue unresolved. These results caution against reading surface metric gains as evidence of true multimodal integration, with direct implications for clinical VLM deployment.

临床AI视觉语言模型幻觉检测神经影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。