评估13种方法在胸部X光多标签分类中的不确定性量化效果
Benchmarking Uncertainty and its Disentanglement in multi-label Chest X-Ray Classification
- 用信息论方法评估神经网络对预测、认知和随机不确定性的量化能力
- 发现不同模型架构与方法在区分认知与随机不确定性上表现差异明显
- 首次将三种不确定性方法扩展至多标签医学图像任务,适用于临床决策支持
可靠的不确定性量化对医疗影像中可信决策和AI模型部署至关重要。尽管已有研究在自然图像分类等理想数据集上探索了神经网络对预测、认知和随机不确定性的信息论量化方法,但其在真实医疗诊断任务中的适用性仍待深入。本研究基于MIMIC-CXR-JPG数据集,为多标签胸部X光分类任务构建了全面的不确定性量化基准。评估了13种不确定性量化方法在卷积神经网络(ResNet)与视觉变换器(Vision Transformer)架构上的表现。同时,将证据深度学习、异类分类网络(HetClass NNs)和确定性不确定性模型扩展至多标签场景。分析揭示了不同方法与架构在不确定性估计有效性及认知/随机不确定性解耦能力上的优劣,为临床可解释性提供依据。
原文摘要 · Abstract (English)
Reliable uncertainty quantification is crucial for trustworthy decision-making and the deployment of AI models in medical imaging. While prior work has explored the ability of neural networks to quantify predictive, epistemic, and aleatoric uncertainties using an information-theoretical approach in synthetic or well defined data settings like natural image classification, its applicability to real life medical diagnosis tasks remains underexplored. In this study, we provide an extensive uncertainty quantification benchmark for multi-label chest X-ray classification using the MIMIC-CXR-JPG dataset. We evaluate 13 uncertainty quantification methods for convolutional (ResNet) and transformer-based (Vision Transformer) architectures across a wide range of tasks. Additionally, we extend Evidential Deep Learning, HetClass NNs, and Deep Deterministic Uncertainty to the multi-label setting. Our analysis provides insights into uncertainty estimation effectiveness and the ability to disentangle epistemic and aleatoric uncertainties, revealing method- and architecture-specific strengths and limitations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。