提升医疗大模型可信度,精准评估问诊中逐步积累证据下的判断信心。
Towards Reliable Medical LLMs: Benchmarking and Enhancing Confidence Estimation of Large Language Models in Medical Consultation
- 构建首个多轮问诊场景下的医学信心评估基准,引入信息充分性梯度。
- 27种方法对比发现,医疗数据放大了原有信心评估的局限性。
- 提出MedConf框架,通过检索增强生成构建症状画像,可解释地整合信心得分。
大规模语言模型(LLMs)常基于不完整信息做出临床判断,增加误诊风险。现有研究多在单轮静态场景下评估信心,忽视了真实问诊中信心与正确性随证据累积的耦合关系,限制了可靠决策支持。本文提出首个针对真实医疗问诊中多轮交互信心评估的基准,统一三类医学数据用于开放式诊断生成,并引入信息充分性梯度刻画信心-正确性动态变化。我们在该基准上实现并对比27种代表性方法,发现:(1) 医学数据加剧了词元级与一致性级信心方法的固有缺陷;(2) 医学推理需同时评估诊断准确性和信息完整性。基于此,我们提出MedConf——一种基于证据的语言自我评估框架,通过检索增强生成构建症状画像,对患者信息进行支持、缺失、矛盾关系对齐,并经加权集成生成可解释的信心估计。在两个LLM和三个医学数据集上,MedConf在AUROC与皮尔逊相关系数上均持续优于现有方法,在信息不足及多重疾病条件下保持稳定性能。结果表明,信息充分性是可信医疗信心建模的关键决定因素,为构建更可靠、可解释的大规模医疗模型提供了新路径。
原文摘要 · Abstract (English)
Large-scale language models (LLMs) often offer clinical judgments based on incomplete information, increasing the risk of misdiagnosis. Existing studies have primarily evaluated confidence in single-turn, static settings, overlooking the coupling between confidence and correctness as clinical evidence accumulates during real consultations, which limits their support for reliable decision-making. We propose the first benchmark for assessing confidence in multi-turn interaction during realistic medical consultations. Our benchmark unifies three types of medical data for open-ended diagnostic generation and introduces an information sufficiency gradient to characterize the confidence-correctness dynamics as evidence increases. We implement and compare 27 representative methods on this benchmark; two key insights emerge: (1) medical data amplifies the inherent limitations of token-level and consistency-level confidence methods, and (2) medical reasoning must be evaluated for both diagnostic accuracy and information completeness. Based on these insights, we present MedConf, an evidence-grounded linguistic self-assessment framework that constructs symptom profiles via retrieval-augmented generation, aligns patient information with supporting, missing, and contradictory relations, and aggregates them into an interpretable confidence estimate through weighted integration. Across two LLMs and three medical datasets, MedConf consistently outperforms state-of-the-art methods on both AUROC and Pearson correlation coefficient metrics, maintaining stable performance under conditions of information insufficiency and multimorbidity. These results demonstrate that information adequacy is a key determinant of credible medical confidence modeling, providing a new pathway toward building more reliable and interpretable large medical models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。