让大模型学会说‘我不确定’,提升医疗AI的可信与安全
The challenge of uncertainty quantification of large language models in medicine
- 融合贝叶斯、集成学习等方法量化认知与随机不确定性
- 通过不确定性地图和置信度指标增强临床可解释性
- 适合医疗AI研发者与关注可信决策的临床应用者
本研究探讨大语言模型在医疗应用中不确定性量化的问题,强调技术突破与哲学意义。随着大模型日益参与临床决策,准确表达不确定性对保障AI辅助医疗的可靠性、安全性与伦理性至关重要。我们提出将不确定性视为知识的必要组成部分,倡导动态反思的AI设计思路。通过结合贝叶斯推断、深度集成、蒙特卡洛丢弃等概率方法,以及预测熵与语义熵的语义分析,构建涵盖认知与随机不确定性的综合框架。引入代理建模克服专有API限制,多源数据融合增强上下文理解,通过持续学习与元学习实现动态校准。通过不确定性图与置信度指标嵌入可解释性,支持用户信任与临床可读性。该方法符合负责任与反思性AI原则。哲学上主张接受可控模糊,而非追求绝对确定,承认医学知识的固有暂定性。
原文摘要 · Abstract (English)
This study investigates uncertainty quantification in large language models (LLMs) for medical applications, emphasizing both technical innovations and philosophical implications. As LLMs become integral to clinical decision-making, accurately communicating uncertainty is crucial for ensuring reliable, safe, and ethical AI-assisted healthcare. Our research frames uncertainty not as a barrier but as an essential part of knowledge that invites a dynamic and reflective approach to AI design. By integrating advanced probabilistic methods such as Bayesian inference, deep ensembles, and Monte Carlo dropout with linguistic analysis that computes predictive and semantic entropy, we propose a comprehensive framework that manages both epistemic and aleatoric uncertainties. The framework incorporates surrogate modeling to address limitations of proprietary APIs, multi-source data integration for better context, and dynamic calibration via continual and meta-learning. Explainability is embedded through uncertainty maps and confidence metrics to support user trust and clinical interpretability. Our approach supports transparent and ethical decision-making aligned with Responsible and Reflective AI principles. Philosophically, we advocate accepting controlled ambiguity instead of striving for absolute predictability, recognizing the inherent provisionality of medical knowledge.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。