提出SynSeg,通过特征协同提升多类别对比学习的分割精度。
SynSeg: Feature Synergy for Multi-Category Contrastive Learning in End-to-End Open-Vocabulary Semantic Segmentation
- 设计特征协同结构,融合先验与激活图重建判别特征
- 在多个基准上实现0.6%至8.9%的mIoU提升
- 轻量级端到端架构,适合实时语义分割场景
开放词汇语义分割面临类别多样性和粒度差异的挑战。现有弱监督方法常依赖类别特定监督和不合适的对比学习特征构建,导致语义错位与性能下降。本文提出一种新弱监督方法SynSeg,通过多类别对比学习(MCCL)注入类内与类间知识,增强训练信号。提出特征协同结构(FSS),通过先验融合与语义激活图增强重构判别特征,有效缓解视觉编码器引入的前景偏差。SynSeg为轻量级端到端方案,支持实时推理。大量实验表明,该方法在所有基准上均超越现有最先进水平,mIoU提升达0.6%至8.9%。
原文摘要 · Abstract (English)
Semantic segmentation in open-vocabulary scenarios presents significant challenges due to the wide range and granularity of semantic categories. Existing weakly-supervised methods often rely on category-specific supervision and ill-suited feature construction methods for contrastive learning, leading to semantic misalignment and poor performance. In this work, we introduce a novel weakly-supervised approach, SynSeg, to address the challenges. SynSeg performs Multi-Category Contrastive Learning (MCCL) as a stronger training signal which robustly injecting intra- and inter-category knowledge during training. We also propose a new feature reconstruction framework named Feature Synergy Structure (FSS). FSS reconstructs discriminative features for contrastive learning through prior fusion and semantic-activation-map enhancement, effectively avoiding the foreground bias introduced by the visual encoder. Furthermore, SynSeg is a lightweight end-to-end solution capable for real-time inference. In general, SynSeg effectively improves the abilities in semantic localization and discrimination under weak supervision in an efficient manner. Extensive experiments on benchmarks demonstrate that our method outperforms state-of-the-art (SOTA) performance, with mIoU score gains ranging from 0.6% up to 8.9% across all reported benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。