arXiv:2607.26886cs.CVcs.AI2026-07

大模型在无图像时会根据患者信息编造诊断,且结果有系统性偏差。

Hearsay: Vision-Language Medical Diagnoses Without an Image

论文配图:Hearsay: Vision-Language Medical Diagnoses Without an Image
图 1 · 摘自论文原文
  • 仅输入患者人口学特征,大模型即生成具体疾病诊断。
  • 65岁白人男性被频繁误诊为黑色素瘤,年轻黑人女性常被诊断为结节病。
  • 需直接审计结构化输出,关注提示词敏感性以评估可靠性。

当要求描述一张从未提供的医学影像时,前沿视觉语言模型不会拒绝回答,而是产生虚构的诊断。我们发现这种虚构并非随机,而是受患者身份影响。在胸部X光、脑部MRI和皮肤科图像中,对Claude Opus-4.7、GPT-5.4和Gemini-3.1-Pro仅提供人口统计描述而无图像,改变描述会系统性地改变返回的诊断结果。Claude表现出高度集中:一名65岁白人男性询问皮肤痣时,几乎每次都给出黑色素瘤诊断;一名32岁黑人女性询问胸片时,被诊断为结节病,其推理显示“基于人口学特征和典型模式推测”。GPT-5.4的影响更广泛,在所有测试的人口学组中均虚构诊断,尤其在年轻黑人患者胸部影像上显著提及结节病。两个结构性发现凸显问题本质:存在一种“含糊状态”,即文本承认图像缺失,但结构化诊断字段仍命名具体疾病,此现象无法通过纯文本审核发现;此外,当将‘skin mole’替换为‘skin lesion’时,Claude的皮肤科错误完全消失,而GPT-5.4的错误仍存,表明此类幻觉是多种不同失效模式的集合,而非单一现象。临床流程中部署可信视觉语言模型,需直接审计结构化输出通道,提示词敏感性应作为首要评估维度。

原文摘要 · Abstract (English)

When asked to describe a medical image that was never attached, frontier vision-language models do not abstain: they confabulate a diagnosis. We show that this confabulation is not random. It is structured by who the patient is said to be. Across chest X-ray, brain MRI, and dermatology, Claude Opus-4.7, GPT-5.4, and Gemini-3.1-Pro are each queried with only a demographic descriptor and no image, and changing the descriptor systematically shifts the diagnosis returned. Claude concentrates sharply: a 65-year-old white man asking about a skin mole receives Melanoma in nearly every response, and a 32-year-old Black woman asking about her chest X-ray receives a Sarcoidosis diagnosis whose reasoning reads "suspected, based on demographics and classic pattern.'' GPT-5.4's effect is broader, fabricating across every demographic cell we test, most conspicuously naming Sarcoidosis for young Black patients on chest X-ray. Two structural findings sharpen the problem. A hedged regime appears in which the prose acknowledges the missing image while the structured diagnosis field nevertheless names a disease, a dissociation invisible to prose-only audits. And Claude's dermatology effect collapses entirely when 'skin mole' is swapped for 'skin lesion' while GPT-5.4's is preserved, indicating that mirage is a family of distinct failure modes rather than a single phenomenon. Trustworthy VLM deployment in clinical pipelines requires auditing the structured output channel directly, and probe-word sensitivity should be treated as a first-class evaluation dimension

医疗AI幻觉检测模型偏见结构化输出

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。