arXiv:2602.06674cs.CVcs.HC2026-02

构建首个病理图像多专家标注基准数据集,助力真实医疗影像分析模型研发。

CytoCrowd: A Multi-Annotator Benchmark Dataset for Cytology Image Analysis

  • 四名病理医生对446张高分辨率图像独立标注,呈现真实专家分歧。
  • 额外提供资深专家制定的高质量金标准,支持客观评估与训练。
  • 适用于检测、分类及标注聚合算法测试,推动下一代医学影像模型发展。

高质量标注数据集对推动医学图像分析中的机器学习至关重要。然而当前存在关键缺口:多数数据集仅提供单一清晰真值,掩盖了真实世界中专家间的分歧;或虽有多重标注,却缺乏独立金标准以进行客观评估。为此,我们提出CytoCrowd,一个面向细胞学分析的新公开基准数据集。该数据集包含446张高分辨率图像,每张图像包含两部分:(1) 四名独立病理医生提供的原始、相互冲突的标注;(2) 由资深专家建立的独立高质量金标准。这一双结构设计使CytoCrowd兼具多重用途:既可作为标准计算机视觉任务(如目标检测与分类)的基准,使用金标准进行评估;同时亦可作为真实场景下评估标注聚合算法的有效性平台。我们提供了两项任务的完整基线结果。实验揭示了该数据集带来的挑战,确立其在开发下一代医学图像分析模型中的重要价值。

原文摘要 · Abstract (English)

High-quality annotated datasets are crucial for advancing machine learning in medical image analysis. However, a critical gap exists: most datasets either offer a single, clean ground truth, which hides real-world expert disagreement, or they provide multiple annotations without a separate gold standard for objective evaluation. To bridge this gap, we introduce CytoCrowd, a new public benchmark for cytology analysis. The dataset features 446 high-resolution images, each with two key components: (1) raw, conflicting annotations from four independent pathologists, and (2) a separate, high-quality gold-standard ground truth established by a senior expert. This dual structure makes CytoCrowd a versatile resource. It serves as a benchmark for standard computer vision tasks, such as object detection and classification, using the ground truth. Simultaneously, it provides a realistic testbed for evaluating annotation aggregation algorithms that must resolve expert disagreements. We provide comprehensive baseline results for both tasks. Our experiments demonstrate the challenges presented by CytoCrowd and establish its value as a resource for developing the next generation of models for medical image analysis.

医学图像标注数据多专家基准数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。