用梯度方差检测超声心动图标注错误,训练中自动修复提升模型鲁棒性
Detecting and refurbishing ground truth errors during training of deep learning-based echocardiography segmentation models

- 通过梯度方差分析识别训练中的错误标注
- 在50%系统性错误下,模型仍保持较好性能
- 提出伪标注修复策略,高错误率时显著提效
基于深度学习的医学图像分割通常依赖人工标注的真值(GT)标签,但这些标签可能包含随机错误或系统性偏差。本研究评估了深度学习模型在超声心动图分割任务中对这类错误的鲁棒性,并提出一种在训练过程中检测并修复错误标签的新策略。使用CAMUS数据集,模拟三类错误,对比基于损失的标注错误检测方法与基于梯度方差(VOG)的方法。同时提出一种伪标注方法用于修复疑似错误的真值标签。在不同错误水平下评估所提方法性能。结果显示,VOG在训练期间能有效识别错误标注;标准U-Net在随机错误和中等系统性错误(最高达50%)下仍表现稳健;所提检测与修复方法在高错误条件下显著提升模型性能。
原文摘要 · Abstract (English)
Deep learning-based medical image segmentation typically relies on ground truth (GT) labels obtained through manual annotation, but these can be prone to random errors or systematic biases. This study examines the robustness of deep learning models to such errors in echocardiography (echo) segmentation and evaluates a novel strategy for detecting and refurbishing erroneous labels during model training. Using the CAMUS dataset, we simulate three error types, then compare a loss-based GT label error detection method with one based on Variance of Gradients (VOG). We also propose a pseudo-labelling approach to refurbish suspected erroneous GT labels. We assess the performance of our proposed approach under varying error levels. Results show that VOG proved highly effective in flagging erroneous GT labels during training. However, a standard U-Net maintained strong performance under random label errors and moderate levels of systematic errors (up to 50%). The detection and refurbishment approach improved performance, particularly under high-error conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。