arXiv:2508.06146cs.CV2025-08ICCV被引 4

用文本引导视觉提示,实现更准的通用图像分割。

Text-guided Visual Prompt DINO for Generic Segmentation

  • 早期融合文本与视觉提示,增强跨模态交互。
  • 对齐查询顺序提升语义空间一致性,准确率显著提高。
  • 自动生成5亿条带标注数据,降噪超80%,适合开放场景研究者。

近期多模态视觉模型在开放世界分割中面临晚期特征融合、混合提示查询选择不佳及依赖描述词库的局限。为此,我们提出Prompt-DINO框架,包含三项创新:首先,在初始编码阶段实现文本/视觉提示与主干特征的早期融合,深化跨模态交互以解决语义歧义;其次,为DETR架构设计顺序对齐的查询选择机制,在解码过程中显式优化文本与视觉查询的结构对齐,提升语义-空间一致性;第三,基于RAP模型构建生成式数据引擎,通过双路径交叉验证流水线合成0.5亿条多样化训练样本,相比传统方法降低80.5%标签噪声。大量实验表明,Prompt-DINO在开放世界检测基准上达到领先性能,显著扩展了语义覆盖范围,突破固定词汇限制。本工作建立了开放场景下可扩展多模态检测与数据生成的新范式。数据与代码见https://github.com/WeChatCV/WeVisionOne。

原文摘要 · Abstract (English)

Recent advancements in multimodal vision models have highlighted limitations in late-stage feature fusion and suboptimal query selection for hybrid prompts open-world segmentation, alongside constraints from caption-derived vocabularies. To address these challenges, we propose Prompt-DINO, a text-guided visual Prompt DINO framework featuring three key innovations. First, we introduce an early fusion mechanism that unifies text/visual prompts and backbone features at the initial encoding stage, enabling deeper cross-modal interactions to resolve semantic ambiguities. Second, we design order-aligned query selection for DETR-based architectures, explicitly optimizing the structural alignment between text and visual queries during decoding to enhance semantic-spatial consistency. Third, we develop a generative data engine powered by the Recognize Anything via Prompting (RAP) model, which synthesizes 0.5B diverse training instances through a dual-path cross-verification pipeline, reducing label noise by 80.5% compared to conventional approaches. Extensive experiments demonstrate that Prompt-DINO achieves state-of-the-art performance on open-world detection benchmarks while significantly expanding semantic coverage beyond fixed-vocabulary constraints. Our work establishes a new paradigm for scalable multimodal detection and data generation in open-world scenarios. Data&Code are available at https://github.com/WeChatCV/WeVisionOne.

通用分割多模态提示学习数据生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。