arXiv:2603.17520cs.CV2026-03中稿 · CVPR被引 5

提出并行聚合框架,提升开放词汇语义与部件分割精度

PCA-Seg: Revisiting Cost Aggregation for Open-Vocabulary Semantic and Part Segmentation

  • 设计并行结构分离语义与空间信息聚合
  • 每块仅增0.35M参数,性能达当前最优
  • 适合需要高精度分割的视觉语言任务研究者

视觉语言模型(VLMs)在开放词汇语义与部件分割(OSPS)中取得显著进展。然而,现有方法通过串行结构依次进行空间与类别聚合,导致类别语义与空间上下文间知识干扰。为此,本文提出简单高效的并行成本聚合(PCA-Seg)范式,缓解该问题,使模型能从成本体积中捕获更丰富的视觉-语言对齐信息。具体地,设计专家驱动感知学习(EPL)模块,高效融合语义与上下文流;引入多专家解析器,从多视角提取互补特征;设计系数映射器,自适应学习每个像素的特征权重,实现互补知识的统一融合。此外,提出特征正交解耦(FOD)策略,降低语义与上下文流间的冗余,使EPL模块能从正交化特征中学习多样化知识。在八个基准上的大量实验表明,每个并行模块仅增加0.35M参数,即达到当前最优的OSPS性能。

原文摘要 · Abstract (English)

Recent advances in vision-language models (VLMs) have garnered substantial attention in open-vocabulary semantic and part segmentation (OSPS). However, existing methods extract image-text alignment cues from cost volumes through a serial structure of spatial and class aggregations, leading to knowledge interference between class-level semantics and spatial context. Therefore, this paper proposes a simple yet effective parallel cost aggregation (PCA-Seg) paradigm to alleviate the above challenge, enabling the model to capture richer vision-language alignment information from cost volumes. Specifically, we design an expert-driven perceptual learning (EPL) module that efficiently integrates semantic and contextual streams. It incorporates a multi-expert parser to extract complementary features from multiple perspectives. In addition, a coefficient mapper is designed to adaptively learn pixel-specific weights for each feature, enabling the integration of complementary knowledge into a unified and robust feature embedding. Furthermore, we propose a feature orthogonalization decoupling (FOD) strategy to mitigate redundancy between the semantic and contextual streams, which allows the EPL module to learn diverse knowledge from orthogonalized features. Extensive experiments on eight benchmarks show that each parallel block in PCA-Seg adds merely 0.35M parameters while achieving state-of-the-art OSPS performance.

语义分割视觉语言并行结构开放词汇

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。