用户用自然语言指导,精准识别物体细节部件。
Fine-Grained Open-Vocabulary Object Recognition via User-Guided Segmentation
- 基于视觉基础模型,通过自然语言引导实现细粒度检测。
- 在多种物体类型上表现稳定,显著提升基线模型能力。
- 用户干预少但控制力强,适合需要精确定制的场景。
近期视觉基础模型的发展使得高效高质量的目标检测成为可能。尽管已有研究取得成功,现有目标检测模型仍难以捕捉整体物体中的微小组件,且无法有效体现用户意图。为此,我们提出一种新型基于基础模型的检测方法FOCUS:通过用户引导分割实现细粒度开放词汇目标识别。FOCUS融合视觉基础模型能力,实现灵活粒度的开放词汇目标检测,并允许用户通过自然语言直接引导检测过程。它不仅能准确识别和定位细微组成元素,还能在减少不必要的用户操作的同时赋予用户显著控制权。借助FOCUS,用户可发出可解释的指令,主动引导检测朝预期方向进行。实验结果表明,FOCUS有效提升了基线模型的检测能力,并在不同物体类型上表现出一致性能。
原文摘要 · Abstract (English)
Recent advent of vision-based foundation models has enabled efficient and high-quality object detection at ease. Despite the success of previous studies, object detection models face limitations on capturing small components from holistic objects and taking user intention into account. To address these challenges, we propose a novel foundation model-based detection method called FOCUS: Fine-grained Open-Vocabulary Object ReCognition via User-Guided Segmentation. FOCUS merges the capabilities of vision foundation models to automate open-vocabulary object detection at flexible granularity and allow users to directly guide the detection process via natural language. It not only excels at identifying and locating granular constituent elements but also minimizes unnecessary user intervention yet grants them significant control. With FOCUS, users can make explainable requests to actively guide the detection process in the intended direction. Our results show that FOCUS effectively enhances the detection capabilities of baseline models and shows consistent performance across varying object types.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。