医学大模型常产生误导性回答,影响诊疗决策。
Medical Hallucinations in Foundation Models and Their Impact on Healthcare
- 通过链式思考提示降低幻觉,提升判断准确性
- 通用模型幻觉率低于专科模型,达76.6%无错率
- 医生审计发现多数错误源于因果推理失误
基础模型中的幻觉源于自回归训练目标对词元似然的过度优化,导致过度自信与不确定性校准不足。我们将医学幻觉定义为任何事实错误、逻辑不一致或缺乏权威临床证据支持、可能改变临床决策的生成内容。我们在七项涵盖医学推理与生物医学信息检索的任务上评估了11个基础模型(7个通用型,4个医学专用型)。通用模型幻觉率显著低于医学专用模型(中位数:76.6% vs 51.3%,差异=25.2%,95%置信区间:18.7-31.3%,Mann-Whitney U = 27.0,p = 0.012,秩相关系数r = -0.64)。顶尖模型Gemini-2.5 Pro在引入链式思考提示后准确率超97%(基础为87.6%),而医学专用模型MedGemma准确率仅为28.6%-61.9%,即便经过医学语料专门训练。链式思考在86.4%的测试对比中显著减少幻觉(经FDR校正,q < 0.05),表明显式推理路径可实现自我验证与错误检测。医师审核确认,64-72%的残留幻觉源于因果或时间推理失败,而非知识缺失。全球临床医生调查(n = 70)验证其真实影响:91.8%曾遭遇医学幻觉,84.7%认为其可能造成患者伤害。医学专用模型表现不佳表明,安全性来自大规模预训练中发展出的复杂推理能力与广义知识整合,而非窄域优化。
原文摘要 · Abstract (English)
Hallucinations in foundation models arise from autoregressive training objectives that prioritize token-likelihood optimization over epistemic accuracy, fostering overconfidence and poorly calibrated uncertainty. We define medical hallucination as any model-generated output that is factually incorrect, logically inconsistent, or unsupported by authoritative clinical evidence in ways that could alter clinical decisions. We evaluated 11 foundation models (7 general-purpose, 4 medical-specialized) across seven medical hallucination tasks spanning medical reasoning and biomedical information retrieval. General-purpose models achieved significantly higher proportions of hallucination-free responses than medical-specialized models (median: 76.6% vs 51.3%, difference = 25.2%, 95% CI: 18.7-31.3%, Mann-Whitney U = 27.0, p = 0.012, rank-biserial r = -0.64). Top-performing models such as Gemini-2.5 Pro exceeded 97% accuracy when augmented with chain-of-thought prompting (base: 87.6%), while medical-specialized models like MedGemma ranged from 28.6-61.9% despite explicit training on medical corpora. Chain-of-thought reasoning significantly reduced hallucinations in 86.4% of tested comparisons after FDR correction (q < 0.05), demonstrating that explicit reasoning traces enable self-verification and error detection. Physician audits confirmed that 64-72% of residual hallucinations stemmed from causal or temporal reasoning failures rather than knowledge gaps. A global survey of clinicians (n = 70) validated real-world impact: 91.8% had encountered medical hallucinations, and 84.7% considered them capable of causing patient harm. The underperformance of medical-specialized models despite domain training indicates that safety emerges from sophisticated reasoning capabilities and broad knowledge integration developed during large-scale pre-training, not from narrow optimization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。