arXiv:2512.09315cs.CV2025-12被引 5

构建医疗影像噪声标签基准,评估10种方法在真实噪声下的表现

Benchmarking Real-World Medical Image Classification with Noisy Labels: Challenges, Practice, and Outlook

  • 构建涵盖7个数据集、6种模态的噪声标签评估框架
  • 发现现有方法在高噪声下性能显著下降,尤其受类别不平衡影响
  • 开源代码库,助力开发更抗噪的医疗影像模型

从噪声标签中学习仍是医学图像分析中的重大挑战,因标注需专家知识且存在显著的观察者差异,常导致标签不一致或错误。尽管已有大量关于噪声标签学习(LNL)的研究,但现有方法在医学影像中的鲁棒性尚未得到系统评估。为此,我们提出LNMBench,一个面向医学影像噪声标签的综合性基准。该基准包含10种代表性方法,在7个数据集、6种成像模态和3种噪声模式下进行评估,建立了一个统一且可复现的评估框架,用于在真实条件下测试鲁棒性。全面实验表明,现有LNL方法在高噪声和真实噪声场景下性能大幅下降,凸显了医学数据中类别不平衡与域变异性的持续挑战。基于此,我们进一步提出一种简单有效的改进策略,以提升模型在上述条件下的鲁棒性。LNMBench代码库已公开,旨在促进标准化评估、可复现研究,并为科研与实际医疗应用中抗噪声算法的开发提供实践指导。

原文摘要 · Abstract (English)

Learning from noisy labels remains a major challenge in medical image analysis, where annotation demands expert knowledge and substantial inter-observer variability often leads to inconsistent or erroneous labels. Despite extensive research on learning with noisy labels (LNL), the robustness of existing methods in medical imaging has not been systematically assessed. To address this gap, we introduce LNMBench, a comprehensive benchmark for Label Noise in Medical imaging. LNMBench encompasses \textbf{10} representative methods evaluated across 7 datasets, 6 imaging modalities, and 3 noise patterns, establishing a unified and reproducible framework for robustness evaluation under realistic conditions. Comprehensive experiments reveal that the performance of existing LNL methods degrades substantially under high and real-world noise, highlighting the persistent challenges of class imbalance and domain variability in medical data. Motivated by these findings, we further propose a simple yet effective improvement to enhance model robustness under such conditions. The LNMBench codebase is publicly released to facilitate standardized evaluation, promote reproducible research, and provide practical insights for developing noise-resilient algorithms in both research and real-world medical applications.The codebase is publicly available on https://github.com/myyy777/LNMBench.

医学影像噪声标签基准测试鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。