提出新任务与数据集,评估模型在真实世界中同时应对领域和类别变化的能力。
ODOV: Benchmark the Open-Domain Open-Vocabulary Object Detection
- 构建跨15个场景的开放域开放词表检测基准OD-LVIS
- 新方法在46,949张图像上实现1,203类物体的鲁棒检测
- 适合研究开放世界目标检测与视觉语言模型应用的学者
现有研究通常将领域偏移与类别偏移视为独立问题,但在真实场景中二者常同时发生且相互作用,导致检测性能显著下降。为此,我们提出并系统研究一种新问题——开放域开放词表(ODOV)目标检测,旨在评估模型在真实环境中应对复合领域与类别偏移的能力。我们构建了新基准OD-LVIS,包含46,949张跨越15个多样化真实场景的图像,覆盖1,203个类别,用于评估目标检测性能。此外,我们提出一种新型ODOV检测基线,充分借助视觉语言模型(VLM)的多模态对齐能力,引入两项关键机制:域无关类别提示(DAPmt),强化类别语义、弱化领域表示,实现纯类别表征;域投影与嫁接(DP&G)模块,从输入图像中融合域特定特征,使模型能动态适应多样开放领域。这两项组件使模型在真实场景下同时面对类别与领域变化时仍保持有效检测性能。我们对提出的ODOV检测任务进行了广泛基准评估,并报告实验结果。这些结果验证了ODOV任务的合理性、OD-LVIS数据集的实用性以及所提方法的优越性。
原文摘要 · Abstract (English)
Existing studies typically investigate domain shift and category shift as independent problems, however, in real-world scenarios, the two types of shifts often occur simultaneously and interact, leading to significant degradation in detection performance. To address this, we propose and systematically study a novel problem-Open-Domain Open-Vocabulary (ODOV) object detection-which aims to evaluate a model's ability to adapt to the compound domain and category shifts in real-world environments.We construct a new benchmark, OD-LVIS, which contains 46,949 images spanning 15 diverse real-world scenarios and 1,203 categories, for assessing object detection performance. Furthermore, we propose a novel ODOV detection baseline that fully leverages VLM's powerful multi-modal alignment capabilities and introduces two key mechanisms to enhance both category and domain generalization. One is the Domain-Agnostic Category Prompt (DAPmt), which strengthens category semantics while attenuating domain representations, enabling pure category representation. The other is the Domain Projection and Grafting (DP&G) module, which incorporates domain-specific features from input images, allowing the model to dynamically generalize across diverse open domains. These two components enable the model to maintain effective detection performance under simultaneous category and domain variations in real-world scenarios. We provide extensive benchmark evaluations for the proposed ODOV detection task and report experimental results. These results validate the soundness of the ODOV task, the practicality of the OD-LVIS dataset, and the superiority of the method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。