arXiv:2506.23751cs.CV2025-06被引 1

用生成图像挑战开放词汇目标检测模型,发现其对位置敏感的漏洞。

Can We Challenge Open-Vocabulary Object Detectors with Generated Content in Street Scenes?

  • 用Stable Diffusion生成多样语义的异常物体进行图像修复
  • 模型在合成数据上普遍漏检,且依赖物体位置而非语义
  • 适合研究模型泛化能力或安全应用评估的学者

开放词汇目标检测器(如Grounding DINO)在大规模多样数据上训练,表现优异,但其局限性尚不明确,尤其在高安全性场景中令人担忧。真实世界数据难以提供可控的评估条件。本文通过合成数据系统探索模型边界,设计两个自动化管道,利用Stable Diffusion结合WordNet和ChatGPT生成高语义多样性的异常物体进行图像修复。基于LostAndFound与NuImages两个真实数据集构建的合成数据,评估多个开放词汇检测器及经典检测器。结果表明,该方法能有效挑战模型,导致大量漏检;且模型对物体位置高度敏感,而非语义。这为系统性测试开放词汇模型提供了新方法,并揭示了改进数据采集的关键方向。

原文摘要 · Abstract (English)

Open-vocabulary object detectors such as Grounding DINO are trained on vast and diverse data, achieving remarkable performance on challenging datasets. Due to that, it is unclear where to find their limitations, which is of major concern when using in safety-critical applications. Real-world data does not provide sufficient control, required for a rigorous evaluation of model generalization. In contrast, synthetically generated data allows to systematically explore the boundaries of model competence/generalization. In this work, we address two research questions: 1) Can we challenge open-vocabulary object detectors with generated image content? 2) Can we find systematic failure modes of those models? To address these questions, we design two automated pipelines using stable diffusion to inpaint unusual objects with high diversity in semantics, by sampling multiple substantives from WordNet and ChatGPT. On the synthetically generated data, we evaluate and compare multiple open-vocabulary object detectors as well as a classical object detector. The synthetic data is derived from two real-world datasets, namely LostAndFound, a challenging out-of-distribution (OOD) detection benchmark, and the NuImages dataset. Our results indicate that inpainting can challenge open-vocabulary object detectors in terms of overlooking objects. Additionally, we find a strong dependence of open-vocabulary models on object location, rather than on object semantics. This provides a systematic approach to challenge open-vocabulary models and gives valuable insights on how data could be acquired to effectively improve these models.

目标检测生成内容模型评估合成数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。