arXiv:2506.16058cs.CV2025-06ICCV被引 6

新基准挑战开放词汇分割,发现旧方法在真实场景下表现大幅下降

Stepping Out of Similar Semantic Space for Open-Vocabulary Segmentation

  • 构建新基准OpenBench,打破训练与测试语义重叠
  • 旧模型在新基准上性能显著下降,暴露能力缺陷
  • 提出OVSNet融合异构特征,实现零成本训练空间扩展

开放词汇分割旨在仅凭无限文本提示实现任意类别分割。现有工作虽借助大规模视觉语言模型取得进展,但其测试集语义空间与训练集高度重合,难以真实评估模型泛化能力。为此,我们提出新基准OpenBench,其语义分布显著区别于训练数据。在该基准上测试现有方法,发现其性能与原有结论严重背离。同时提出OVSNet方法,通过精细融合异构特征并零成本扩展训练空间,在现有数据集和OpenBench上均达到当前最优结果。分析验证了基准与方法的有效性。

原文摘要 · Abstract (English)

Open-vocabulary segmentation aims to achieve segmentation of arbitrary categories given unlimited text inputs as guidance. To achieve this, recent works have focused on developing various technical routes to exploit the potential of large-scale pre-trained vision-language models and have made significant progress on existing benchmarks. However, we find that existing test sets are limited in measuring the models' comprehension of ``open-vocabulary" concepts, as their semantic space closely resembles the training space, even with many overlapping categories. To this end, we present a new benchmark named OpenBench that differs significantly from the training semantics. It is designed to better assess the model's ability to understand and segment a wide range of real-world concepts. When testing existing methods on OpenBench, we find that their performance diverges from the conclusions drawn on existing test sets. In addition, we propose a method named OVSNet to improve the segmentation performance for diverse and open scenarios. Through elaborate fusion of heterogeneous features and cost-free expansion of the training space, OVSNet achieves state-of-the-art results on both existing datasets and our proposed OpenBench. Corresponding analysis demonstrate the soundness and effectiveness of our proposed benchmark and method.

开放词汇图像分割基准测试视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。