arXiv:2604.14651cs.CL2026-04ACL

让医疗大模型的预测更靠谱,误差越大的病人越不敢乱下结论。

CURA: Clinical Uncertainty Risk Alignment for Language Model-Based Risk Prediction

论文配图:CURA: Clinical Uncertainty Risk Alignment for Language Model-Based Risk Prediction
图 1 · 摘自论文原文
  • 用双层级目标优化模型,同时对个体和群体不确定性建模。
  • 在MIMIC-IV数据上,校准度提升但判别力基本不变。
  • 适合需要可信风险评估的临床决策支持系统使用。

临床语言模型(LM)越来越多地用于从自由文本病历中进行风险预测,但其不确定性估计常未校准且临床不可靠。本文提出临床不确定性风险对齐(CURA)框架,将临床LM的风险预测与个体错误概率及群体层面的模糊性对齐。CURA首先微调领域特定的临床LM以获得任务适配的患者嵌入表示,再通过双层不确定性目标对多头分类器进行不确定性微调。具体而言,个体级校准项使预测不确定性与每位患者的出错概率对齐;群体感知正则项则将风险估计拉向嵌入空间局部邻域中的事件率,并在决策边界附近的模糊群体上施加更大权重。进一步分析表明,该正则项可解释为基于邻域信息的软标签交叉熵损失,提供了标签平滑视角。在MIMIC-IV多个临床风险预测任务上的大量实验显示,CURA持续提升校准性能,且未显著损害判别能力。进一步分析表明,CURA降低了过度自信的误判风险,生成了更可信的不确定性估计,适用于下游临床决策支持。

原文摘要 · Abstract (English)

Clinical language models (LMs) are increasingly applied to support clinical risk prediction from free-text notes, yet their uncertainty estimates often remain poorly calibrated and clinically unreliable. In this work, we propose Clinical Uncertainty Risk Alignment (CURA), a framework that aligns clinical LM-based risk estimates and uncertainty with both individual error likelihoods and cohort-level ambiguities. CURA first fine-tunes domain-specific clinical LMs to obtain task-adapted patient embeddings, and then performs uncertainty fine-tuning of a multi-head classifier using a bi-level uncertainty objective. Specifically, an individual-level calibration term aligns predictive uncertainty with each patient's likelihood of error, while a cohort-aware regularizer pulls risk estimates toward event rates in their local neighborhoods in the embedding space and places extra weight on ambiguous cohorts near the decision boundary. We further show that this cohort-aware term can be interpreted as a cross-entropy loss with neighborhood-informed soft labels, providing a label-smoothing view of our method. Extensive experiments on MIMIC-IV clinical risk prediction tasks across various clinical LMs show that CURA consistently improves calibration metrics without substantially compromising discrimination. Further analysis illustrates that CURA reduces overconfident false reassurance and yields more trustworthy uncertainty estimates for downstream clinical decision support.

医疗AI不确定性估计风险预测大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。