提出X-Agent框架,让模型更好理解未见类别的语义特征。
Novel Category Discovery with X-Agent Attention for Open-Vocabulary Semantic Segmentation
- 用可感知语义的'代理'控制跨模态注意力,动态优化隐式语义
- 在多个基准上达到最新性能,显著提升未见类别的语义清晰度
- 适合研究开放词汇语义分割和视觉语言模型解释性的学者
开放词汇语义分割(OVSS)通过文本驱动对齐实现像素级分类,但基础类别训练与开放词汇推理间的领域差异,导致对隐式未见类别的区分建模困难。现有基于视觉-语言模型(VLM)的方法虽借助预训练多模态表示取得良好效果,但其隐式语义理解机制仍不明确,成为制约OVSS的关键瓶颈。本文首次开展探测实验,分析在归纳学习范式下VLM中隐式语义的分布模式与动态特性。基于此,提出X-Agent框架,引入感知隐式语义的‘代理’,协同调控跨模态注意力机制,同步优化隐式语义动态并增强其可感知性。大量基准测试表明,X-Agent实现当前最优性能,并有效提升隐式语义显著性。
原文摘要 · Abstract (English)
Open-vocabulary semantic segmentation (OVSS) conducts pixel-level classification via text-driven alignment, where the domain discrepancy between base category training and open-vocabulary inference poses challenges in discriminative modeling of latent unseen category. To address this challenge, existing vision-language model (VLM)-based approaches demonstrate commendable performance through pre-trained multi-modal representations. However, the fundamental mechanisms of latent semantic comprehension remain underexplored, making the bottleneck for OVSS. In this work, we initiate a probing experiment to explore distribution patterns and dynamics of latent semantics in VLMs under inductive learning paradigms. Building on these insights, we propose X-Agent, an innovative OVSS framework employing latent semantic-aware ``agent'' to orchestrate cross-modal attention mechanisms, simultaneously optimizing latent semantic dynamic and amplifying its perceptibility. Extensive benchmark evaluations demonstrate that X-Agent achieves state-of-the-art performance while effectively enhancing the latent semantic saliency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。