arXiv:2410.14365cs.CV2024-10被引 5

研究病理图像标注噪声对CNN模型训练的影响,提出防过拟合有效策略。

Impact of imperfect annotations on CNN training and performance for instance segmentation and classification in digital pathology

  • 用小规模正确标注验证集控制噪声过拟合
  • 预训练显著提升模型在噪声数据下的性能
  • 适合医学图像分析与弱标注场景的研究者

在数字病理学中,对大量实例(如细胞核)进行分割与分类是准确诊断的关键任务。然而,由于标注过程复杂,高质量深度学习数据集常难以获取。本文研究了噪声标注对最先进的CNN模型在组织病理图像中联合检测、分割和分类任务训练与性能的影响。我们探讨了确定合适训练轮数的条件,以防止模型过拟合到标注噪声。结果表明,使用小规模且正确标注的验证集可有效避免过拟合,并在很大程度上维持模型性能。此外,研究还强调了预训练的积极作用。

原文摘要 · Abstract (English)

Segmentation and classification of large numbers of instances, such as cell nuclei, are crucial tasks in digital pathology for accurate diagnosis. However, the availability of high-quality datasets for deep learning methods is often limited due to the complexity of the annotation process. In this work, we investigate the impact of noisy annotations on the training and performance of a state-of-the-art CNN model for the combined task of detecting, segmenting and classifying nuclei in histopathology images. In this context, we investigate the conditions for determining an appropriate number of training epochs to prevent overfitting to annotation noise during training. Our results indicate that the utilisation of a small, correctly annotated validation set is instrumental in avoiding overfitting and maintaining model performance to a large extent. Additionally, our findings underscore the beneficial role of pre-training.

医学图像弱监督标注噪声深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。