arXiv:2503.00359cs.CV2025-03CVPR被引 10

针对开放世界实例检测,提出新方法提升匹配精度。

Solving Instance Detection from an Open-World Perspective

  • 利用开放世界数据微调预训练模型,优化实例级特征匹配
  • 通过新颖视角合成与干扰项采样,增强视觉参考多样性
  • 在两个新基准上显著超越此前方法,性能提升超10个AP

实例检测(InsDet)旨在基于给定视觉参考,在新场景图像中定位特定物体实例。技术上需先生成候选区域,再进行实例级匹配以识别目标。其开放世界特性支持机器人、AR/VR等广泛应用,但也带来挑战:测试场景图像未在训练中出现,且视觉参考与检测候选间存在领域差异。现有方法通过合成训练样本或使用现成基础模型应对,但未能充分利用开放世界信息。本文从开放世界视角出发,提出方法IDOW。发现预训练基础模型虽召回率高,但不专用于实例级匹配。因此,我们基于开放世界数据微调模型,引入度量学习与新型数据增强策略,包括以干扰项为负样本、合成新视角实例以丰富参考。大量实验表明,本方法在两个最新挑战性基准数据集上,于常规与新实例检测设置下均显著优于先前工作,性能提升超过10 AP。

原文摘要 · Abstract (English)

Instance detection (InsDet) aims to localize specific object instances within a novel scene imagery based on given visual references. Technically, it requires proposal detection to identify all possible object instances, followed by instance-level matching to pinpoint the ones of interest. Its open-world nature supports its broad applications from robotics to AR/VR but also presents significant challenges: methods must generalize to unknown testing data distributions because (1) the testing scene imagery is unseen during training, and (2) there are domain gaps between visual references and detected proposals. Existing methods tackle these challenges by synthesizing diverse training examples or utilizing off-the-shelf foundation models (FMs). However, they only partially capitalize the available open-world information. In contrast, we approach InsDet from an Open-World perspective, introducing our method IDOW. We find that, while pretrained FMs yield high recall in instance detection, they are not specifically optimized for instance-level feature matching. Therefore, we adapt pretrained FMs for improved instance-level matching using open-world data. Our approach incorporates metric learning along with novel data augmentations, which sample distractors as negative examples and synthesize novel-view instances to enrich the visual references. Extensive experiments demonstrate that our method significantly outperforms prior works, achieving >10 AP over previous results on two recently released challenging benchmark datasets in both conventional and novel instance detection settings.

实例检测开放世界度量学习视觉匹配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。