医学标注不确定性强,忽略它会误导模型评估结果。
Clinical Uncertainty Impacts Machine Learning Evaluations
- 用概率度量直接处理标签置信度分布,而非简单投票
- 在医疗图像数据集上,考虑置信度后模型排名显著变化
- 方法轻量高效,适合公开原始标注供复现
临床数据标注很少确定,标注者意见不一且置信度不均。传统聚合方法如多数投票会掩盖这种变异性。在医学影像基准上的简单实验表明,考虑二值标签的置信度会显著影响模型排名。因此我们主张机器学习评估应显式使用基于概率的度量来处理标注不确定性,这些度量可独立于标注生成过程(如计数、主观评分或概率响应模型)应用。它们计算轻量,一旦按模型得分排序,即可线性时间实现闭式表达。我们呼吁社区公开数据集原始标注,并采用考虑不确定性的评估方式,使性能估计更真实反映临床数据特征。
原文摘要 · Abstract (English)
Clinical dataset labels are rarely certain as annotators disagree and confidence is not uniform across cases. Typical aggregation procedures, such as majority voting, obscure this variability. In simple experiments on medical imaging benchmarks, accounting for the confidence in binary labels significantly impacts model rankings. We therefore argue that machine-learning evaluations should explicitly account for annotation uncertainty using probabilistic metrics that directly operate on distributions. These metrics can be applied independently of the annotations' generating process, whether modeled by simple counting, subjective confidence ratings, or probabilistic response models. They are also computationally lightweight, as closed-form expressions have linear-time implementations once examples are sorted by model score. We thus urge the community to release raw annotations for datasets and to adopt uncertainty-aware evaluation so that performance estimates may better reflect clinical data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。