arXiv:2412.00076cs.CV2024-12被引 14

剖析ImageNet-1k数据集的标签错误与设计缺陷,推动计算机视觉基准改进。

Flaws of ImageNet, Computer Vision's Favourite Dataset

  • 分析图像标签错误、类别定义模糊等核心问题
  • 发现训练与评估域存在显著偏差,且存在重复图像
  • 呼吁学术界共同优化这一关键基准数据集

自发布以来,ImageNet-1k已成为评估模型性能的黄金标准,并成为众多计算机视觉数据集和训练任务的基础。随着模型准确率提升,标签正确性等问题日益凸显。本文分析了ImageNet-1k中的多项缺陷,包括错误标签、类别定义重叠或模糊、训练-评估域偏移以及图像重复。部分问题解决方案明确,但更多问题亟需学界展开深入讨论,以进一步完善这一极具影响力的基准数据集,更好地支持未来研究。

原文摘要 · Abstract (English)

Since its release, ImageNet-1k dataset has become a gold standard for evaluating model performance. It has served as the foundation for numerous other datasets and training tasks in computer vision. As models have improved in accuracy, issues related to label correctness have become increasingly apparent. In this blog post, we analyze the issues in the ImageNet-1k dataset, including incorrect labels, overlapping or ambiguous class definitions, training-evaluation domain shifts, and image duplicates. The solutions for some problems are straightforward. For others, we hope to start a broader conversation about refining this influential dataset to better serve future research.

ImageNet数据质量计算机视觉

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。