测试大模型诊断可靠性,发现它们易受无关信息干扰且判断常不靠谱。
The Reliability of LLMs for Medical Diagnosis: An Examination of Consistency, Manipulation, and Contextual Awareness
- 通过改写、添加无关细节和补充临床信息,评估模型表现
- 两款模型一致性100%,但分别有40%和30%被无关信息误导
- 医生认为模型对上下文的反应中,谷歌模型更准确,但仍需人工把关
本研究评估了两款大语言模型(Google Gemini 2.0 Flash 和 OpenAI ChatGPT-4o)在三个维度上的诊断可靠性:对重述输入的一致性、对无关提示内容的敏感性,以及对附加临床信息的响应能力。设计了52个临床场景,并在控制条件下进行修改。一致性测试通过改变人口学特征、表述方式和检查项目来实现,保持诊断核心不变;敏感性测试在不改变临床证据的前提下,加入看似合理但无关的叙述细节;上下文感知测试则增加患者病史、生活方式数据或诊断结果以引导预期诊断。由医生评审判断上下文引发的诊断变化是否临床上合理。结果显示,两款模型在所有等效变体和重复查询中均保持100%一致性。当加入无关细节时,Gemini 在40.0% 的案例中更改诊断,ChatGPT 在30.0% 中更改。尽管 ChatGPT 更频繁响应上下文(77.8% 对 55.6%),但其变更中临床不合理比例更高(33.3% 对 22.2%)。相比之下,Gemini 的上下文驱动调整更常被认定为恰当(66.7% 对 55.6%)。一致性并未阻止模型在输入被操纵或上下文变化时出现非合理的诊断偏移。在大模型能有效区分相关信息与噪音、并识别证据不足之前,其诊断应用必须依赖医生监督与结构化防护机制。
原文摘要 · Abstract (English)
This study evaluated the diagnostic reliability of two Large Language Models (LLMs), Google Gemini 2.0 Flash and OpenAI ChatGPT-4o, across three dimensions: consistency under rephrased inputs, susceptibility to irrelevant prompt content, and responsiveness to added clinical context. We designed 52 clinical scenarios and modified each under controlled conditions. For consistency, scenarios were rephrased with demographic, wording, and examination changes that preserved the diagnostic core. And the susceptibility was evaluated through embedding irrelevant but plausible narrative details while keeping the clinical evidence unchanged. For contextual awareness, patient history, lifestyle data, or diagnostic findings were added to shift the expected diagnosis. Physician reviewers then judged whether context-driven changes were clinically appropriate. Both models returned identical diagnoses across all equivalent variants and repeated queries (100% consistency). When irrelevant details were added, Gemini changed its diagnosis in 40.0% of cases and ChatGPT in 30.0%. ChatGPT responded to context more often than Gemini (77.8% vs. 55.6%), but a larger share of its changes were clinically inappropriate (33.3% vs. 22.2%). Gemini's context-driven changes were more often judged appropriate (66.7% vs. 55.6%). Consistency under controlled inputs did not protect either model from irrelevant manipulation or unjustified diagnostic shifts when context changed. Before LLMs can separate relevant from irrelevant input and flag insufficient evidence, their diagnostic use requires clinician oversight and structured safeguards.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。