arXiv:2602.06938cs.CVcs.LG2026-02中稿 · IEEE Engineering i…

提出医疗影像误标注检测框架,提升胶囊内镜数据质量

Reliable Mislabel Detection for Video Capsule Endoscopy Data

  • 基于深度学习构建误标注检测流程,自动识别异常标签
  • 在两大公开胶囊内镜数据集上验证,清理后异常检测性能提升
  • 经三位专家复核确认,适合医学影像数据清洗场景

深度神经网络的分类性能高度依赖大规模、精准标注的数据集。但在医疗影像领域,由于标注需由专业医师完成,标注者资源有限,且类别边界常模糊难定义,导致高质量数据获取困难。本文针对此问题,提出一种医学数据集中的误标注检测框架,验证于两个最大的公开胶囊内镜数据集。通过该流程识别出潜在误标样本,并由三位资深胃肠科医生进行复审与重新标注。结果表明,该框架能有效检测错误标签,数据清洗后异常检测性能优于现有基线方法。

原文摘要 · Abstract (English)

The classification performance of deep neural networks relies strongly on access to large, accurately annotated datasets. In medical imaging, however, obtaining such datasets is particularly challenging since annotations must be provided by specialized physicians, which severely limits the pool of annotators. Furthermore, class boundaries can often be ambiguous or difficult to define which further complicates machine learning-based classification. In this paper, we want to address this problem and introduce a framework for mislabel detection in medical datasets. This is validated on the two largest, publicly available datasets for Video Capsule Endoscopy, an important imaging procedure for examining the gastrointestinal tract based on a video stream of lowresolution images. In addition, potentially mislabeled samples identified by our pipeline were reviewed and re-annotated by three experienced gastroenterologists. Our results show that the proposed framework successfully detects incorrectly labeled data and results in an improved anomaly detection performance after cleaning the datasets compared to current baselines.

医疗影像误标注检测胶囊内镜

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。