用医生标注置信度提升超声肺部图像分割精度,改善临床决策。
Label Uncertainty for Ultrasound Segmentation
- 让医生标注每个像素的置信度,建模真实医疗数据中的不确定性。
- 以60%置信度阈值筛选标签训练,分割性能显著优于传统方法。
- 提升分割质量后,临床任务如氧合比预测、再入院风险评估更准确。
在医学影像中,放射科医生之间的观察差异常导致标注不确定性,尤其在主观性较强的模态中。肺部超声(LUS)是典型例子——其图像常混合高度模糊区域与清晰结构,即使经验丰富的医师也难以一致标注。本文提出一种新方法,通过专家提供的逐像素置信度进行标注与模型训练。不将标注视为绝对真值,而是设计标注协议捕捉医生对各区域判断的信心,建模真实临床数据中的固有随机不确定性。实验表明,将置信度融入训练可提升分割性能;更重要的是,这种提升带来下游临床任务的改进:包括估算S/F氧合比、分类S/F比变化及预测30天再入院。我们评估了多种引入不确定性的方法,发现使用60%置信度阈值二值化标签进行训练效果最佳,远优于50%阈值的朴素做法,说明仅用高置信像素训练更有效。本研究系统分析不同置信阈值的影响,不仅比较分割指标,还评估临床结果。结果表明,标签置信度是重要信号,合理利用可显著提升AI在医学影像中的可靠性与临床实用性。
原文摘要 · Abstract (English)
In medical imaging, inter-observer variability among radiologists often introduces label uncertainty, particularly in modalities where visual interpretation is subjective. Lung ultrasound (LUS) is a prime example-it frequently presents a mixture of highly ambiguous regions and clearly discernible structures, making consistent annotation challenging even for experienced clinicians. In this work, we introduce a novel approach to both labeling and training AI models using expert-supplied, per-pixel confidence values. Rather than treating annotations as absolute ground truth, we design a data annotation protocol that captures the confidence that radiologists have in each labeled region, modeling the inherent aleatoric uncertainty present in real-world clinical data. We demonstrate that incorporating these confidence values during training leads to improved segmentation performance. More importantly, we show that this enhanced segmentation quality translates into better performance on downstream clinically-critical tasks-specifically, estimating S/F oxygenation ratio values, classifying S/F ratio change, and predicting 30-day patient readmission. While we empirically evaluate many methods for exposing the uncertainty to the learning model, we find that a simple approach that trains a model on binarized labels obtained with a (60%) confidence threshold works well. Importantly, high thresholds work far better than a naive approach of a 50% threshold, indicating that training on very confident pixels is far more effective. Our study systematically investigates the impact of training with varying confidence thresholds, comparing not only segmentation metrics but also downstream clinical outcomes. These results suggest that label confidence is a valuable signal that, when properly leveraged, can significantly enhance the reliability and clinical utility of AI in medical imaging.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。