arXiv:2607.20582cs.LGcs.AI2026-07

用不确定性估计提升医疗AI决策可靠性

Bayesian uncertainty estimation improves clinical decision making in medical AI agents

论文配图:Bayesian uncertainty estimation improves clinical decision making in medical AI agents
图 1 · 摘自论文原文
  • 用蒙特卡洛丢弃法获取模型对诊断的置信度信号
  • 将误判检测准确率提升至AUROC 0.77,降低错误率56%
  • 适合需要可解释性医疗辅助系统的开发者

医学图像分析中的机器学习模型通常缺乏可靠的置信度衡量,限制了其在模糊或非典型病例中的应用。本文在包含8种胸部异常、137,593张训练图像的多任务胸片分类器上,采用蒙特卡洛丢弃法,获得了一种表征认知不确定性的信号。该信号能追踪模型在不同训练集规模下的泛化能力,并标记出看似自信但易出错的预测。将此信号加入点预测后,误判检测的AUROC从0.74提升至0.77(ΔAUROC +0.023,95% CI [+0.014, +0.033])。在一项2×2受控实验中,临床决策支持系统仅在接收二值错误风险标志而非原始分数时有效利用该信息,将不可靠发现上的自信误诊率从8.5%降至2.7%。表明认知不确定性蕴含超越点预测的决策价值,但其实际效用取决于传递方式。

原文摘要 · Abstract (English)

Machine learning models for medical image analysis typically lack a reliable measure of confidence, limiting their use in ambiguous or atypical cases. Here we show that Monte Carlo dropout, applied to a multi-task chest-radiograph classifier (eight thoracic findings, 137,593 training images), provides an epistemic uncertainty signal that tracks generalisation across training-set scales and flags confident yet error-prone predictions. Adding this signal to the point prediction raised error-detection AUROC from 0.74 to 0.77 ($Δ$AUROC +0.023, 95% CI [+0.014, +0.033]). In a controlled 2x2 factorial experiment, a clinical-decision-support agent exploited this uncertainty only when it was delivered as a binary error-risk flag rather than as raw scores, cutting confident misdiagnoses on unreliable findings from 8.5% to 2.7%. Epistemic uncertainty estimation thus carries decision-relevant information beyond point predictions, but its value for downstream agents depends on how it is communicated.

医疗AI不确定性估计决策支持

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。