arXiv:2505.20928cs.CV2025-05中稿 · MICCAI 2026

研究发现:医学图像预训练无需高精度标注,可节省大量人工成本。

Good Enough? An Investigation on the Impact of Label Quality in Large-Scale Medical Datasets

  • 对比不同质量标签对预训练模型的影响
  • 高质量标签对预训练无显著提升作用
  • 适合医疗数据预训练与资源优化研究者

放射科分割标注的手动精修成本极高。为评估其在模型训练中的必要性,我们研究了标签质量与模型性能的关系。突破以往仅关注直接推理模型的局限,首次系统分析预训练数据集中的标签质量影响。结果表明,尽管部署前模型仍需高质量标签,但预训练阶段对标签精度要求并不严格。该发现质疑了大规模数据集进行全人工精修的必要性,建议将专家精力集中于下游目标任务的高质量数据构建。

原文摘要 · Abstract (English)

Manually refining radiological segmentation masks is highly resource-intensive. To determine when this expert commitment is truly justified for the training of segmentation models, we investigate the relationship between label quality and model performance. Expanding beyond models trained directly for inference, we conduct the first study isolating the impact of label quality in pre-training datasets. While high-quality labels remain essential for models proceeding directly to deployment, we find no evidence that strict label quality is crucial for pre-training efficacy. These results question the necessity of exhaustive human-in-the-loop refinement for massive corpora intended for pretraining and suggest that expert effort is more effectively invested in well-curated downstream target datasets.

医学图像预训练标签质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。