arXiv:2507.19870cs.CVcs.HC2025-07被引 3

通过人机协作实现小样本开放世界目标检测,提升模型适应新物体能力。

OW-CLIP: Data-Efficient Visual Supervision for Open-World Object Detection via Human-AI Collaboration

  • 引入可插拔的多模态提示调优与裁剪平滑技术,缓解特征过拟合。
  • 仅用3.8%自生成数据达到SOTA 89%性能,且在等量数据下超越现有方法。
  • 提供可视化界面辅助生成高质量标注,适合需要持续更新模型的研究者。

开放世界目标检测(OWOD)需同时识别已知与未知物体,要求模型随新标注持续适应。现有方法存在三大问题:依赖大量众包标注、易受部分特征过拟合影响、需修改模型结构灵活性差。为此,我们提出OW-CLIP,一个视觉分析系统,支持数据高效增量训练。该系统采用专为OWOD设计的即插即用多模态提示调优,并提出新颖的“裁剪平滑”技术以缓解过拟合。为满足训练数据需求,我们提出双模态数据精炼方法,利用大语言模型与跨模态相似性进行数据生成与过滤。同时开发可视化界面,支持用户探索并生成类别特定的视觉特征短语与细粒度差异化图像。定量评估表明,OW-CLIP仅需3.8%自生成数据即可达到SOTA 89%的性能;在相同数据量下优于当前最优方法。案例研究验证了方法有效性及标注质量提升。

原文摘要 · Abstract (English)

Open-world object detection (OWOD) extends traditional object detection to identifying both known and unknown object, necessitating continuous model adaptation as new annotations emerge. Current approaches face significant limitations: 1) data-hungry training due to reliance on a large number of crowdsourced annotations, 2) susceptibility to "partial feature overfitting," and 3) limited flexibility due to required model architecture modifications. To tackle these issues, we present OW-CLIP, a visual analytics system that provides curated data and enables data-efficient OWOD model incremental training. OW-CLIP implements plug-and-play multimodal prompt tuning tailored for OWOD settings and introduces a novel "Crop-Smoothing" technique to mitigate partial feature overfitting. To meet the data requirements for the training methodology, we propose dual-modal data refinement methods that leverage large language models and cross-modal similarity for data generation and filtering. Simultaneously, we develope a visualization interface that enables users to explore and deliver high-quality annotations: including class-specific visual feature phrases and fine-grained differentiated images. Quantitative evaluation demonstrates that OW-CLIP achieves competitive performance at 89% of state-of-the-art performance while requiring only 3.8% self-generated data, while outperforming SOTA approach when trained with equivalent data volumes. A case study shows the effectiveness of the developed method and the improved annotation quality of our visualization system.

目标检测开放世界人机协作小样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。