arXiv:2509.01275cs.CV2025-09中稿 · ACMMM2025被引 3

提出X-Agent框架,让模型更好理解未见类别的语义特征。

Novel Category Discovery with X-Agent Attention for Open-Vocabulary Semantic Segmentation

  • 用可感知语义的'代理'控制跨模态注意力,动态优化隐式语义
  • 在多个基准上达到最新性能,显著提升未见类别的语义清晰度
  • 适合研究开放词汇语义分割和视觉语言模型解释性的学者

开放词汇语义分割(OVSS)通过文本驱动对齐实现像素级分类,但基础类别训练与开放词汇推理间的领域差异,导致对隐式未见类别的区分建模困难。现有基于视觉-语言模型(VLM)的方法虽借助预训练多模态表示取得良好效果,但其隐式语义理解机制仍不明确,成为制约OVSS的关键瓶颈。本文首次开展探测实验,分析在归纳学习范式下VLM中隐式语义的分布模式与动态特性。基于此,提出X-Agent框架,引入感知隐式语义的‘代理’,协同调控跨模态注意力机制,同步优化隐式语义动态并增强其可感知性。大量基准测试表明,X-Agent实现当前最优性能,并有效提升隐式语义显著性。

原文摘要 · Abstract (English)

Open-vocabulary semantic segmentation (OVSS) conducts pixel-level classification via text-driven alignment, where the domain discrepancy between base category training and open-vocabulary inference poses challenges in discriminative modeling of latent unseen category. To address this challenge, existing vision-language model (VLM)-based approaches demonstrate commendable performance through pre-trained multi-modal representations. However, the fundamental mechanisms of latent semantic comprehension remain underexplored, making the bottleneck for OVSS. In this work, we initiate a probing experiment to explore distribution patterns and dynamics of latent semantics in VLMs under inductive learning paradigms. Building on these insights, we propose X-Agent, an innovative OVSS framework employing latent semantic-aware ``agent'' to orchestrate cross-modal attention mechanisms, simultaneously optimizing latent semantic dynamic and amplifying its perceptibility. Extensive benchmark evaluations demonstrate that X-Agent achieves state-of-the-art performance while effectively enhancing the latent semantic saliency.

语义分割视觉语言模型开放词汇

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。