arXiv:2505.03910cs.CL2025-05

探究医学影像模型的预测不确定性与医生报告中语言不确定性的关联

Hesitation is defeat? Connecting Linguistic and Predictive Uncertainty

  • 用BERT分析放射科报告,提取人类语言不确定性标签
  • 蒙特卡洛丢弃与深度集成在预测不确定性上表现良好
  • 机器与人类不确定性相关性较弱,需进一步优化临床适配

利用深度学习自动化胸部X光片解读有望显著提升临床流程、决策效率和大规模健康筛查。然而,在医疗场景中,仅优化预测性能不足,量化不确定性同样关键。本文研究了基于贝叶斯深度学习近似得出的预测不确定性,与由规则化标注器标记的自由文本放射科报告中的人类/语言不确定性之间的关系。选用BERT作为模型,评估了不同二值化不确定性标签的方法,并探讨了蒙特卡洛丢弃和深度集成在预测不确定性估计中的效果。结果表明模型性能良好,但预测不确定性与语言不确定性之间相关性仅为中等水平,凸显了机器不确定性与人类判断细微差别对齐的挑战。研究提示,尽管贝叶斯近似能提供有价值的不确定性估计,但在临床应用中仍需进一步改进以充分捕捉和利用人类不确定性特征。

原文摘要 · Abstract (English)

Automating chest radiograph interpretation using Deep Learning (DL) models has the potential to significantly improve clinical workflows, decision-making, and large-scale health screening. However, in medical settings, merely optimising predictive performance is insufficient, as the quantification of uncertainty is equally crucial. This paper investigates the relationship between predictive uncertainty, derived from Bayesian Deep Learning approximations, and human/linguistic uncertainty, as estimated from free-text radiology reports labelled by rule-based labellers. Utilising BERT as the model of choice, this study evaluates different binarisation methods for uncertainty labels and explores the efficacy of Monte Carlo Dropout and Deep Ensembles in estimating predictive uncertainty. The results demonstrate good model performance, but also a modest correlation between predictive and linguistic uncertainty, highlighting the challenges in aligning machine uncertainty with human interpretation nuances. Our findings suggest that while Bayesian approximations provide valuable uncertainty estimates, further refinement is necessary to fully capture and utilise the subtleties of human uncertainty in clinical applications.

医学影像不确定性贝叶斯深度学习BERT

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。