arXiv:2503.14862cs.CV2025-03ICRA被引 2

提出细粒度开放词汇检测新任务与数据集,解决评估不公问题。

Fine-Grained Open-Vocabulary Object Detection with Fined-Grained Prompts: Task, Dataset and Benchmark

  • 设计3F-OVD任务,要求模型理解细粒度描述与图像细节。
  • 构建17.1万标注样本的NEU-171K数据集,支持监督与开放词汇场景。
  • 提供新后处理方法,提升模型在细粒度物体上的检测精度。

开放词汇检测旨在定位并识别新类别物体。然而,视觉语言词汇数据的差异会导致评估不公平且不可靠。现有方法尝试通过引入物体属性或位置特征来改进,但这些信息依赖图像具体细节,需人工精确描述,限制了模型性能。本文提出3F-OVD新任务,将细粒度监督检测扩展至开放词汇设置,要求模型深入理解细粒度描述并关注图像中的细微特征。针对高质量细粒度数据集稀缺问题,构建了包含17.1万样本的NEU-171K数据集,覆盖监督与开放词汇两种训练范式。在该数据集上对主流检测器进行基准测试,并提出一种简单有效的后处理技术。代码、数据与标注已开源于https://github.com/tengerye/3FOVD。

原文摘要 · Abstract (English)

Open-vocabulary detectors are proposed to locate and recognize objects in novel classes. However, variations in vision-aware language vocabulary data used for open-vocabulary learning can lead to unfair and unreliable evaluations. Recent evaluation methods have attempted to address this issue by incorporating object properties or adding locations and characteristics to the captions. Nevertheless, since these properties and locations depend on the specific details of the images instead of classes, detectors can not make accurate predictions without precise descriptions provided through human annotation. This paper introduces 3F-OVD, a novel task that extends supervised fine-grained object detection to the open-vocabulary setting. Our task is intuitive and challenging, requiring a deep understanding of Fine-grained captions and careful attention to Fine-grained details in images in order to accurately detect Fine-grained objects. Additionally, due to the scarcity of qualified fine-grained object detection datasets, we have created a new dataset, NEU-171K, tailored for both supervised and open-vocabulary settings. We benchmark state-of-the-art object detectors on our dataset for both settings. Furthermore, we propose a simple yet effective post-processing technique. Our data, annotations and codes are available at https://github.com/tengerye/3FOVD.

开放词汇细粒度检测数据集目标检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。