arXiv:2604.01898cs.LG2026-04

用专家意见提升医疗AI的不确定性估计,让系统更可靠。

Enhancing the Reliability of Medical AI through Expert-guided Uncertainty Modeling

  • 利用专家分歧生成训练目标,分离估算数据噪声和模型不确定性的成分。
  • 在多种任务中提升不确定性估计质量9%至50%。
  • 适合构建高风险医疗AI系统,尤其关注可靠性与可解释性。

人工智能系统在医疗领域加速工作流程并提高诊断准确性,常作为第二意见工具。然而,AI错误的不可预测性带来重大挑战,尤其在医疗场景中可能造成严重后果。目前普遍采用为预测结果附加不确定性估计的方式,使专家能聚焦高风险案例,简化常规验证。但现有不确定性估计方法仍有限,尤其难以量化由数据模糊性和噪声引起的随机不确定性(aleatoric uncertainty)。为此,我们提出一种新方法:利用专家回答的不一致生成训练目标,结合标准数据标签,通过两组模型架构及其实用轻量版本,分别估计总方差中的两个分量。我们在二分类图像识别、二分类与多分类图像分割以及选择题问答任务上验证该方法。实验表明,融入专家知识可使不确定性估计质量提升9%至50%,具体取决于任务类型,证明这一信息源对构建风险感知型医疗AI系统至关重要。

原文摘要 · Abstract (English)

Artificial intelligence (AI) systems accelerate medical workflows and improve diagnostic accuracy in healthcare, serving as second-opinion systems. However, the unpredictability of AI errors poses a significant challenge, particularly in healthcare contexts, where mistakes can have severe consequences. A widely adopted safeguard is to pair predictions with uncertainty estimation, enabling human experts to focus on high-risk cases while streamlining routine verification. Current uncertainty estimation methods, however, remain limited, particularly in quantifying aleatoric uncertainty, which arises from data ambiguity and noise. To address this, we propose a novel approach that leverages disagreement in expert responses to generate targets for training machine learning models. These targets are used in conjunction with standard data labels to estimate two components of uncertainty separately, as given by the law of total variance, via a two-ensemble approach, as well as its lightweight variant. We validate our method on binary image classification, binary and multi-class image segmentation, and multiple-choice question answering. Our experiments demonstrate that incorporating expert knowledge can enhance uncertainty estimation quality by $9\%$ to $50\%$ depending on the task, making this source of information invaluable for the construction of risk-aware AI systems in healthcare applications.

医疗AI不确定性估计专家知识风险感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。