解决多类别语义分割中概念冲突问题,提升模型稳定性。
CoCo-SAM3: Harnessing Concept Conflict in Open-Vocabulary Semantic Segmentation

- 分离类内增强与类间竞争,统一不同类别判断标准。
- 无需训练,在8个基准上实现稳定性能提升。
- 有效缓解同义词导致的类别混淆和重叠问题。
SAM3通过提示驱动的掩码生成范式推进了开放词汇语义分割。但在多类别开放词汇场景中,从不同类别提示独立生成的掩码缺乏统一且可比的证据尺度,常导致覆盖重叠和不稳定的类间竞争。此外,同一概念的不同同义表达倾向于激活不一致的语义与空间证据,引发类内漂移,加剧类间冲突,影响整体推理稳定性。为此,我们提出CoCo-SAM3(Concept-Conflict SAM3),显式地将推理过程解耦为类内增强与类间竞争。方法首先对同义提示的证据进行对齐与聚合,强化概念一致性;随后在统一可比尺度上进行类间竞争,支持所有候选类别间的像素级直接比较。该机制稳定了多类别推理,有效缓解类间冲突。无需额外训练,CoCo-SAM3在八个开放词汇语义分割基准上均实现一致性能提升。
原文摘要 · Abstract (English)
SAM3 advances open-vocabulary semantic segmentation by introducing a prompt-driven mask generation paradigm. However, in multi-class open-vocabulary scenarios, masks generated independently from different category prompts lack a unified and inter-class comparable evidence scale, often resulting in overlapping coverage and unstable competition. Moreover, synonymous expressions of the same concept tend to activate inconsistent semantic and spatial evidence, leading to intra-class drift that exacerbates inter-class conflicts and compromises overall inference stability. To address these issues, we propose CoCo-SAM3 (Concept-Conflict SAM3), which explicitly decouples inference into intra-class enhancement and inter-class competition. Our method first aligns and aggregates evidence from synonymous prompts to strengthen concept consistency. It then performs inter-class competition on a unified comparable scale, enabling direct pixel-wise comparisons among all candidate classes. This mechanism stabilizes multi-class inference and effectively mitigates inter-class conflicts. Without requiring any additional training, CoCo-SAM3 achieves consistent improvements across eight open-vocabulary semantic segmentation benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。