通过词汇对齐提升无源域适应的分割精度
Lost in Translation? Vocabulary Alignment for Source-Free Adaptation in Open-Vocabulary Semantic Segmentation
- 用教师-学生框架+词汇对齐,生成更准伪标签
- 在CityScapes上提升6.11 mIoU,零样本性能领先
- 轻量适配结合选顶类机制,适合资源受限场景
我们提出VocAlign,一种专为视觉语言模型在开放词汇语义分割中设计的无源域适应框架。该方法采用教师-学生范式,并引入词汇对齐策略,通过融入额外类别概念提升伪标签生成质量。为保证效率,使用低秩适配(LoRA)微调模型,保留原始能力的同时降低计算开销。此外,提出针对学生模型的Top-K类别选择机制,显著减少内存占用并进一步提升适应性能。该方法在CityScapes数据集上实现6.11 mIoU的显著提升,在零样本分割基准测试中表现优异,为开放词汇设置下的无源域适应树立了新标准。
原文摘要 · Abstract (English)
We introduce VocAlign, a novel source-free domain adaptation framework specifically designed for VLMs in open-vocabulary semantic segmentation. Our method adopts a student-teacher paradigm enhanced with a vocabulary alignment strategy, which improves pseudo-label generation by incorporating additional class concepts. To ensure efficiency, we use Low-Rank Adaptation (LoRA) to fine-tune the model, preserving its original capabilities while minimizing computational overhead. In addition, we propose a Top-K class selection mechanism for the student model, which significantly reduces memory requirements while further improving adaptation performance. Our approach achieves a notable 6.11 mIoU improvement on the CityScapes dataset and demonstrates superior performance on zero-shot segmentation benchmarks, setting a new standard for source-free adaptation in the open-vocabulary setting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。