为提升复杂物体检测精度,构建含560万属性标注的新数据集
An Attribute-Enriched Dataset and Auto-Annotated Pipeline for Open Detection
- 在Objects365基础上添加颜色、材质等5类属性标注
- 包含140万框级标注,共560万条属性描述
- 可用于改进语言引导的开放目标检测模型
通过语言检测目标常面临挑战,尤其当目标不常见或描述复杂时,自动化模型与人工标注间存在感知差异。为此,我们提出Objects365-Attr数据集,作为现有Objects365数据集的扩展,其核心特征是引入了详尽的属性标注。该数据集通过整合颜色、材质、状态、纹理和色调等多类属性,显著降低物体检测中的不一致性。全数据集包含560万条对象级属性描述,覆盖140万边界框。为验证数据集有效性,我们对YOLO-World在不同规模下的检测性能进行了严格评估,结果表明该数据集能有效推动开放目标检测的发展。
原文摘要 · Abstract (English)
Detecting objects of interest through language often presents challenges, particularly with objects that are uncommon or complex to describe, due to perceptual discrepancies between automated models and human annotators. These challenges highlight the need for comprehensive datasets that go beyond standard object labels by incorporating detailed attribute descriptions. To address this need, we introduce the Objects365-Attr dataset, an extension of the existing Objects365 dataset, distinguished by its attribute annotations. This dataset reduces inconsistencies in object detection by integrating a broad spectrum of attributes, including color, material, state, texture and tone. It contains an extensive collection of 5.6M object-level attribute descriptions, meticulously annotated across 1.4M bounding boxes. Additionally, to validate the dataset's effectiveness, we conduct a rigorous evaluation of YOLO-World at different scales, measuring their detection performance and demonstrating the dataset's contribution to advancing object detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。