arXiv:2510.10524cs.CV2025-10ICCV被引 8

用多模态提示统一解决开放词汇与上下文分割问题

Unified Open-World Segmentation with Multi-Modal Prompts

  • 通过图文提示联合建模,统一处理多种分割任务
  • 在开放词汇与上下文分割上均显著超越现有方法
  • 图文协同提升泛化能力,适合多模态场景应用

本文提出COSINE,一种统一的开世界分割模型,整合开放词汇分割与上下文分割任务,支持文本和图像等多模态提示。COSINE利用基础模型提取输入图像与多模态提示的表征,并通过SegDecoder对齐表征、建模交互关系,生成不同粒度的掩码。该方法克服了以往两类任务在架构、学习目标与表征策略上的差异。大量实验表明,COSINE在开放词汇分割与上下文分割任务中均有显著性能提升。探索性分析显示,视觉与文本提示的协同作用显著优于单一模态方法,带来更强的泛化能力。

原文摘要 · Abstract (English)

In this work, we present COSINE, a unified open-world segmentation model that consolidates open-vocabulary segmentation and in-context segmentation with multi-modal prompts (e.g., text and image). COSINE exploits foundation models to extract representations for an input image and corresponding multi-modal prompts, and a SegDecoder to align these representations, model their interaction, and obtain masks specified by input prompts across different granularities. In this way, COSINE overcomes architectural discrepancies, divergent learning objectives, and distinct representation learning strategies of previous pipelines for open-vocabulary segmentation and in-context segmentation. Comprehensive experiments demonstrate that COSINE has significant performance improvements in both open-vocabulary and in-context segmentation tasks. Our exploratory analyses highlight that the synergistic collaboration between using visual and textual prompts leads to significantly improved generalization over single-modality approaches.

分割多模态提示学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。