arXiv:2504.15707cs.CVcs.AI2025-04被引 10

修正了POPE基准中图像标注错误,发现标签质量显著影响模型评估结果。

RePOPE: Impact of Annotation Errors on the POPE Benchmark

  • 基于重新标注的RePOPE数据集,评估标注误差对模型表现的影响。
  • 不同子集的标注错误分布不均,导致模型排名发生明显变化。
  • 适合关注基准数据质量与模型评估可靠性的研究者参考。

由于数据标注成本高昂,基准数据集常采用现有图像数据集的标签。本文评估了MSCOCO中标注错误对常用物体幻觉评估基准POPE的影响。我们重新标注了基准图像,发现不同子集间标注错误存在不平衡。在修订后的标签(称为RePOPE)上评估多个模型,观察到模型排名出现显著变化,凸显了标签质量的重要性。代码与数据已公开于https://github.com/YanNeu/RePOPE。

原文摘要 · Abstract (English)

Since data annotation is costly, benchmark datasets often incorporate labels from established image datasets. In this work, we assess the impact of label errors in MSCOCO on the frequently used object hallucination benchmark POPE. We re-annotate the benchmark images and identify an imbalance in annotation errors across different subsets. Evaluating multiple models on the revised labels, which we denote as RePOPE, we observe notable shifts in model rankings, highlighting the impact of label quality. Code and data are available at https://github.com/YanNeu/RePOPE .

模型评估标注误差基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。