用不确定性估计替代学习型拒答,更抗数据分布变化。
Is Uncertainty Quantification a Viable Alternative to Learned Deferral?
- 用不确定性量化判断是否拒答,不依赖特定训练数据。
- 在眼底图像上,不确定度方法在分布外样本中表现更优。
- 适合临床场景中应对真实世界数据漂移的AI系统。
人工智能有望显著提升患者护理水平,但并非绝对可靠,需人机协作确保安全。其中关键一环是模型在可能误判时主动将决策权交给人类专家。现有研究多通过优化代理损失函数来学习何时拒答,但此类方法在临床部署时易受数据分布偏移影响。本文提出假设:基于不确定性量化的拒答策略(无需监督学习)比传统学习型拒答更具鲁棒性。为此,我们在大规模眼科数据集上对比了多种学习型拒答模型与经典不确定性量化方法,在分布内和分布外情形下评估其对青光眼分类的准确率与拒答能力。结果表明,不确定性量化方法在分布外样本中表现更稳定,具有成为可靠拒答机制的潜力。
原文摘要 · Abstract (English)
Artificial Intelligence (AI) holds the potential to dramatically improve patient care. However, it is not infallible, necessitating human-AI-collaboration to ensure safe implementation. One aspect of AI safety is the models' ability to defer decisions to a human expert when they are likely to misclassify autonomously. Recent research has focused on methods that learn to defer by optimising a surrogate loss function that finds the optimal trade-off between predicting a class label or deferring. However, during clinical translation, models often face challenges such as data shift. Uncertainty quantification methods aim to estimate a model's confidence in its predictions. However, they may also be used as a deferral strategy which does not rely on learning from specific training distribution. We hypothesise that models developed to quantify uncertainty are more robust to out-of-distribution (OOD) input than learned deferral models that have been trained in a supervised fashion. To investigate this hypothesis, we constructed an extensive evaluation study on a large ophthalmology dataset, examining both learned deferral models and established uncertainty quantification methods, assessing their performance in- and out-of-distribution. Specifically, we evaluate their ability to accurately classify glaucoma from fundus images while deferring cases with a high likelihood of error. We find that uncertainty quantification methods may be a promising choice for AI deferral.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。