arXiv:2505.11034cs.CVcs.AI2025-05中稿 · Journal of Data-ce…被引 3

首个大规模图像数据清洗基准,助力真实医疗图像质量评估

CleanPatrick: A Benchmark for Image Data Cleaning

  • 基于皮肤科数据集构建,通过众包与专家评审生成高质量标注
  • 发现4%无关样本、21%近似重复、32%标签错误,验证清洗必要性
  • 提供可复现的评估框架,适合医疗视觉模型训练前数据净化

稳健的机器学习依赖于干净的数据,但现有的图像数据清洗基准多依赖合成噪声或狭窄的人类研究,限制了比较性和现实相关性。我们提出 CleanPatrick,首个大规模图像领域数据清洗基准,基于公开的 Fitzpatrick17k 皮肤病数据集构建。我们从 933 名医学众包工作者处收集了 496,377 条二元标注,识别出 4% 的无关样本、21% 的近似重复样本和 32% 的标签错误,并采用受项目反应理论启发的聚合模型结合专家评审,生成高质量真值。CleanPatrick 将问题检测形式化为排序任务,使用符合实际审计流程的标准排序指标。我们在该基准上对经典异常检测器、感知哈希、SSIM、Confident Learning、NoiseRank、FINE、BHN 与 SelfClean 进行评测。结果表明,自监督表征在近似重复检测中表现优异,传统方法在受限审核预算下对无关样本检测具竞争力,而针对细粒度医学分类的不合理标签检测在保守人类判断下仍具挑战。通过发布数据集与评估框架,CleanPatrick 支持图像清洗策略的系统性比较。

原文摘要 · Abstract (English)

Robust machine learning depends on clean data, yet current image data cleaning benchmarks rely on synthetic noise or narrow human studies, limiting comparison and real-world relevance. We introduce CleanPatrick, the first large-scale benchmark for data cleaning in the image domain, built upon the publicly available Fitzpatrick17k dermatology dataset. We collect 496,377 binary annotations from 933 medical crowd workers, identify off-topic samples (4%), near-duplicates (21%), and label errors (32%), and employ an aggregation model inspired by item-response theory followed by expert review to derive high-quality ground truth. CleanPatrick formalizes issue detection as a ranking task and employs standard ranking metrics that mirror real audit workflows. We benchmark classical anomaly detectors, perceptual hashing, SSIM, Confident Learning, NoiseRank, FINE, BHN, and SelfClean. On CleanPatrick, self-supervised representations excel at near-duplicate detection, classical methods achieve competitive off-topic detection under constrained review budgets, and detecting implausible labels under conservative human judgment remains challenging for fine-grained medical classification. By releasing both the dataset and the evaluation framework, CleanPatrick enables a systematic comparison of image-cleaning strategies.

数据清洗医疗图像基准测试众包标注

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。