arXiv:2508.00557cs.CV2025-08ICCV被引 12

无需训练即可提升开放词汇语义分割的类别纯净度

Training-Free Class Purification for Open-Vocabulary Semantic Segmentation

  • 通过净化类别表示,解决类别冗余与语义混淆问题
  • 在8个基准上显著提升分割性能,可即插即用
  • 适合追求高效部署的开放词汇分割研究者

微调预训练视觉语言模型已成为增强开放词汇语义分割(OVSS)的强大方法。然而,在大规模数据集上训练带来的巨大计算和资源开销,促使人们关注无需训练的OVSS方法。现有无需训练的方法主要集中在修改模型结构和生成原型以提升分割性能,但往往忽视了类别冗余(当前测试图像中不存在多个类别)和视觉-语言歧义(类别间语义相似导致激活混淆)带来的挑战。这些问题会导致次优的类别激活图和亲和性优化激活图。受此启发,我们提出FreeCP——一种新颖的无需训练类别净化框架,旨在解决上述问题。FreeCP专注于净化语义类别并纠正由冗余和歧义引发的错误。净化后的类别表示被用于生成最终的分割预测。我们在八个基准上进行了大量实验,验证了FreeCP的有效性。结果表明,FreeCP作为即插即用模块,能显著提升与其他OVSS方法结合时的分割性能。

原文摘要 · Abstract (English)

Fine-tuning pre-trained vision-language models has emerged as a powerful approach for enhancing open-vocabulary semantic segmentation (OVSS). However, the substantial computational and resource demands associated with training on large datasets have prompted interest in training-free methods for OVSS. Existing training-free approaches primarily focus on modifying model architectures and generating prototypes to improve segmentation performance. However, they often neglect the challenges posed by class redundancy, where multiple categories are not present in the current test image, and visual-language ambiguity, where semantic similarities among categories create confusion in class activation. These issues can lead to suboptimal class activation maps and affinity-refined activation maps. Motivated by these observations, we propose FreeCP, a novel training-free class purification framework designed to address these challenges. FreeCP focuses on purifying semantic categories and rectifying errors caused by redundancy and ambiguity. The purified class representations are then leveraged to produce final segmentation predictions. We conduct extensive experiments across eight benchmarks to validate FreeCP's effectiveness. Results demonstrate that FreeCP, as a plug-and-play module, significantly boosts segmentation performance when combined with other OVSS methods.

语义分割视觉语言无训练类别净化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。