arXiv:2412.00153cs.CVcs.LG2024-12被引 6

ROSE实现开放集密集分割,可自动生成新类别并精细优化掩码。

ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model

  • 将图像块作为独立候选区域,支持密集与稀疏掩码同步预测。
  • 无需预设类别提示,在开放集下实现自由生成新类别。
  • 通过对话式迭代优化,提升掩码细节和分类精度,适合复杂场景应用。

CLIP与大型多模态模型(LMM)的发展推动了开放词汇和自由文本分割,但现有模型仍需预定义类别提示,限制了自由生成新类别的能力。多数分割类LMM仅限于稀疏预测,难以适应开放集环境。本文提出ROSE——一种革命性的开放集密集分割多模态模型,通过图像块级感知实现密集掩码预测与开放类别自动生成。方法将每个图像块视为独立兴趣区域候选,实现密集与稀疏掩码的联合预测。同时,设计新型指令-响应范式,充分利用LMM的生成与泛化能力,实现不受封闭集约束的类别预测。为进一步提升掩码细节与类别精度,引入基于对话的迭代优化机制,结合前一步预测结果与文本提示进行修正。大量实验表明,ROSE在统一框架下实现了多种分割任务的竞争力表现。代码将公开。

原文摘要 · Abstract (English)

Advances in CLIP and large multimodal models (LMMs) have enabled open-vocabulary and free-text segmentation, yet existing models still require predefined category prompts, limiting free-form category self-generation. Most segmentation LMMs also remain confined to sparse predictions, restricting their applicability in open-set environments. In contrast, we propose ROSE, a Revolutionary Open-set dense SEgmentation LMM, which enables dense mask prediction and open-category generation through patch-wise perception. Our method treats each image patch as an independent region of interest candidate, enabling the model to predict both dense and sparse masks simultaneously. Additionally, a newly designed instruction-response paradigm takes full advantage of the generation and generalization capabilities of LMMs, achieving category prediction independent of closed-set constraints or predefined categories. To further enhance mask detail and category precision, we introduce a conversation-based refinement paradigm, integrating the prediction result from previous step with textual prompt for revision. Extensive experiments demonstrate that ROSE achieves competitive performance across various segmentation tasks in a unified framework. Code will be released.

密集分割开放集多模态模型自动生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。