arXiv:2411.03497cs.CL2024-11NAACL被引 5

提升医疗语言模型预测的可信度,量化并降低不确定性。

Uncertainty Quantification for Clinical Outcome Predictions with (Large) Language Models

  • 采用多任务与集成方法,在白盒模型中减少预测不确定性。
  • 在超过6000名患者数据上验证,多种场景下不确定性显著下降。
  • 适用于医院部署的开源与闭源大模型,增强临床决策透明性。

为支持医疗交付,语言模型(LMs)在利用电子健康记录(EHRs)进行临床预测方面具有巨大潜力。然而,在高风险应用中,不可靠的决策可能因患者安全受损和伦理问题带来高昂代价,因此亟需对自动化临床预测进行有效的不确定性建模。为此,本文研究了在白盒与黑盒设置下对基于EHR的LMs进行不确定性量化的方法。首先,在可访问模型参数与输出logits的白盒场景中,提出多任务学习与集成方法,有效降低了模型不确定性。在此基础上,将方法扩展至黑盒场景,涵盖GPT-4等主流专有语言模型。基于来自10项临床预测任务、超过6000名患者的纵向临床数据,验证了所提框架的有效性。结果表明,集成方法与多任务预测提示能普遍降低不确定性,提升了白盒与黑盒模型的可解释性,推动了可靠AI在医疗领域的应用。

原文摘要 · Abstract (English)

To facilitate healthcare delivery, language models (LMs) have significant potential for clinical prediction tasks using electronic health records (EHRs). However, in these high-stakes applications, unreliable decisions can result in high costs due to compromised patient safety and ethical concerns, thus increasing the need for good uncertainty modeling of automated clinical predictions. To address this, we consider the uncertainty quantification of LMs for EHR tasks in white- and black-box settings. We first quantify uncertainty in white-box models, where we can access model parameters and output logits. We show that an effective reduction of model uncertainty can be achieved by using the proposed multi-tasking and ensemble methods in EHRs. Continuing with this idea, we extend our approach to black-box settings, including popular proprietary LMs such as GPT-4. We validate our framework using longitudinal clinical data from more than 6,000 patients in ten clinical prediction tasks. Results show that ensembling methods and multi-task prediction prompts reduce uncertainty across different scenarios. These findings increase the transparency of the model in white-box and black-box settings, thus advancing reliable AI healthcare.

医疗AI不确定性量化大模型临床预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。