arXiv:2607.02594cs.CVcs.AI2026-07被引 1

通过分析损失函数序列自动识别医学图像错标,提升模型性能。

An automated method of identifying incorrectly labelled images based on the sequences of loss functions of deep learning networks

  • 用深度网络训练中损失变化序列检测标签错误
  • 在10788张眼底图中检出75.31%的故意错标,误报仅4.85%
  • 适合医疗图像数据清洗,尤其用于糖尿病视网膜病变筛查

深度学习广泛应用于医学图像分析,但高达10%的手动标注图像可能存在错误,降低模型性能。本文提出一种自动化方法,通过分析深度分类网络在多个训练轮次中的损失函数序列,识别错误标注的医学图像。被识别出的图像可交由专家复核并重新标注,从而提升数据集质量与模型性能。在视网膜病变筛查的眼底图像数据集上进行了两项实验验证:第一项中,对10,788个金标准标签中6%(648张)故意调换标签,该方法成功识别出75.31%(488张)的错误样本,正确标签中仅有4.85%(492张)被误判为错误;第二项中,对980个被识别样本(占数据集9.1%)进行复查修正后重新训练模型,独立测试集上的最佳准确率从95.93%(含6%标签噪声)提升至96.50%(含1.5%噪声),接近理想水平96.57%(无噪声)。结果表明,该方法能有效通过自动化标签质量控制提升模型表现。

原文摘要 · Abstract (English)

Deep learning is widely applied in medical image analysis, but up to 10% of manually labelled images may be incorrect, degrading model performance. This paper proposes an automated method to identify incorrectly labelled medical images by analyzing sequences of loss functions from deep learning classification networks over multiple training epochs. Identified images can be reviewed and relabelled by experts, improving dataset quality and model performance. Two experiments validate the method on a fundus image dataset for referable diabetic retinopathy screening. In the first, 6% (648) of 10,788 gold-standard labels were intentionally flipped. The method identified 75.31% (488) of the flipped samples, with only 4.85% (492) false positives among correctly labelled samples. In the second, reviewing and correcting the 980 identified samples (9.1% of the dataset) and retraining the model improved best accuracy on an independent test set from 95.93% (with 6% label noise) to 96.50% (with 1.5% noise), approaching the ideal 96.57% (with 0% noise). The results demonstrate the method's effectiveness in improving model performance through automated label quality control.

医学图像标签纠错深度学习数据质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。