arXiv:2409.09412cs.CVcs.AI2024-09中稿 · WACV 2025, added r…被引 9

提出标签收敛概念,揭示物体识别性能上限受标注矛盾限制。

Label Convergence: Defining an Upper Performance Bound in Object Recognition through Contradictory Annotations

  • 定义标签收敛:在存在矛盾标注下模型性能的理论上限。
  • 实测LVIS数据集标签收敛区间为62.63-67.52 mAP@[0.5:0.95:0.05]。
  • 建议改进评估方式、清理测试数据、引入多标注以暴露标注问题。

标注错误不仅影响模型训练,也影响评估。数据集中标签差异和不准确常表现为违背标注规范的矛盾样本,显著影响均值平均精度(mAP)等指标。本文提出“标签收敛”概念,描述在测试标注存在矛盾时模型可达到的最高性能,即性能上限。基于对包括LVIS在内的五个真实数据集的分析,我们以95%置信度估算出LVIS数据集的标签收敛区间为62.63-67.52 mAP@[0.5:0.95:0.05],归因于实际标注错误。当前最先进(SOTA)模型已处于该区间的上端,表明模型能力足以解决现有目标检测任务。因此未来应聚焦三点:(1) 更新问题定义与评估方法,纳入不可避免的标注噪声;(2) 清理数据,尤其是测试数据;(3) 使用多标注数据,提前发现并量化标注变异。

原文摘要 · Abstract (English)

Annotation errors are a challenge not only during training of machine learning models, but also during their evaluation. Label variations and inaccuracies in datasets often manifest as contradictory examples that deviate from established labeling conventions. Such inconsistencies, when significant, prevent models from achieving optimal performance on metrics such as mean Average Precision (mAP). We introduce the notion of "label convergence" to describe the highest achievable performance under the constraint of contradictory test annotations, essentially defining an upper bound on model accuracy. Recognizing that noise is an inherent characteristic of all data, our study analyzes five real-world datasets, including the LVIS dataset, to investigate the phenomenon of label convergence. We approximate that label convergence is between 62.63-67.52 mAP@[0.5:0.95:0.05] for LVIS with 95% confidence, attributing these bounds to the presence of real annotation errors. With current state-of-the-art (SOTA) models at the upper end of the label convergence interval for the well-studied LVIS dataset, we conclude that model capacity is sufficient to solve current object detection problems. Therefore, future efforts should focus on three key aspects: (1) updating the problem specification and adjusting evaluation practices to account for unavoidable label noise, (2) creating cleaner data, especially test data, and (3) including multi-annotated data to investigate annotation variation and make these issues visible from the outset.

目标检测标注误差性能上限数据质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。