用目标物体特征选图,让标注钱花得更值。
Single-Pass Object-Focused Data Selection
- 基于大模型的物体级特征选数据,覆盖所有目标类别
- 在检测与分割任务中,显著优于随机选图
- 适合预算有限时的主动学习或自动标注
尽管未标注图像数据通常丰富,但高质量标注成本高昂,如何高效利用标注预算成为关键挑战。本文聚焦单次通过数据选择(single-pass data selection),即在训练下游模型前一次性选定所有待标注数据。现有方法依赖图像级表示,在目标检测和分割任务上难以稳定超越随机选择。我们提出物体聚焦数据选择(OFDS),利用基础模型的物体级特征,确保目标类别的语义覆盖。在多个任务与目标域的实验中,OFDS始终优于随机选择与所有基线方法。在标注预算受限时,结合OFDS选出的人工标注与基础模型的自动标注,可取得最佳效果。此外,使用OFDS选取主动学习的初始标注集,能持续提升性能。
原文摘要 · Abstract (English)
While unlabeled image data is often plentiful, the costs of high-quality labels pose an important practical challenge: Which images should one select for labeling to use the annotation budget for a particular target task most effectively? To address this problem, we focus on single-pass data selection, which refers to the process of selecting all data to be annotated at once before training a downstream model. Prior methods for single-pass data selection rely on image-level representations and fail to reliably outperform random selection for object detection and segmentation. We propose Object-Focused Data Selection (OFDS) which leverages object-level features from foundation models and ensures semantic coverage of all target classes. In extensive experiments across tasks and target domains, OFDS consistently outperforms random selection and all baselines. The best results for constrained annotation budgets are obtained by combining human labels from OFDS with autolabels from foundation models. Moreover, using OFDS to select the initial labeled set for active learning yields consistent improvements
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。