arXiv:2605.20110cs.CV2026-05被引 2

用概念集合统一建模多目标指代分割,提升开放场景下的准确性。

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction

论文配图:SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction
图 1 · 摘自论文原文
  • 以自然语言概念作为语义条件,联合解码目标掩码集
  • 在gRefCOCO上提升3.3点gIoU,目标越多优势越明显
  • 支持视频指代分割,7个基准上刷新最佳结果

指代分割将自然语言查询定位到像素级掩码,但面对多实例、跨类别分组或开放目标集等复杂场景仍具挑战。以往基于大视觉语言模型的方法用特殊标记序列表示目标,将多个目标视为独立输出,未能捕捉集合层面的完整性与互斥性。本文提出SetCon,将开放指代分割重新定义为显式的集合级概念预测任务,利用LVLM生成的自然语言概念作为语义条件,进行联合掩码集解码。通过分层语义分解,先预测共享的集合级概念定义目标范围,再细化为与目标子集对齐的细粒度概念组。为此,构建了两阶段标注流程,在现有推理分割数据集上添加分层语义监督(236k样本,784k概念短语)。SetCon在图像基准上达到最优性能(gRefCOCO上+3.3 gIoU,MUSE上+12.1 gIoU),且目标数量越多,优势越显著。该概念接口还成功迁移至视频场景,在七项指代视频基准上取得新最优结果,包括MeViS上+10.9 J&F和Ref-SeCVOS上+12.4 J&F。

原文摘要 · Abstract (English)

Referring segmentation grounds natural-language queries to pixel-level masks, but extending it to complex scenarios with multiple instances, cross-category groups, or open-ended target sets remains challenging. Previous Large Vision Language Model (LVLM)-based methods represent referred targets with one or more special tokens sequentially, treating multiple targets as separate outputs rather than a coherent set and offering little incentive to capture set-level properties such as completeness and mutual exclusivity. We reformulate open-ended referring segmentation as explicit set-level concept prediction and propose Set-Concept Segmentation (SetCon), which uses LVLM-generated natural-language concepts, instead of segmentation-specific tokens, as semantic conditions for joint mask-set decoding. A hierarchical semantic decomposition first predicts a shared set-level concept defining the target scope and then refines it into fine-grained concept groups aligned with target subsets. To support this, a two-stage annotation pipeline augments existing reasoning segmentation datasets with hierarchical semantic supervision (236k samples, 784k concept phrases). SetCon achieves state-of-the-art results on image benchmarks (+3.3 gIoU on gRefCOCO, +12.1 gIoU on MUSE), with margins that grow as the number of referred targets increases. The concept interface also transfers to video under a detect-and-track setting, yielding new state-of-the-art results on seven referring video benchmarks, including +10.9 J&F on MeViS and +12.4 J&F on Ref-SeCVOS.

指代分割集合建模视频理解视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。