用视觉语言模型提升遥感图像半监督分割的伪标签质量
Vision-Language Model Purified Semi-Supervised Semantic Segmentation for Remote Sensing Images
- 引入视觉语言模型净化教师网络生成的伪标签
- 在多类别边界区域显著提升伪标签准确率,实现当前最佳性能
- 方法通用可解释,适合遥感领域半监督学习研究者
半监督语义分割(S4)可从低成本未标注遥感图像中学习丰富视觉知识。然而,传统S4架构普遍面临伪标签质量低的问题,尤其在教师-学生框架中更为突出。本文提出新型SemiEarth模型,引入视觉语言模型(VLMs)解决遥感领域的S4挑战。具体设计了VLM伪标签净化(VLM-PP)结构,显著提升教师网络生成伪标签的质量。尤其在遥感图像多类别边界区域,VLM-PP模块能有效改善伪标签质量,从而正确引导学生模型学习。此外,由于VLM-PP具备开放世界能力且与S4架构解耦,当其预测与伪标签存在不一致时,可自动修正低置信度伪标签中的误分类。我们在多个遥感数据集上进行了广泛实验,结果表明SemiEarth达到当前最优(SOTA)性能。更重要的是,相较于以往的遥感领域S4先进方法,本模型不仅表现优异,还具备良好可解释性。代码已开源:https://github.com/wangshanwen001/SemiEarth。
原文摘要 · Abstract (English)
The semi-supervised semantic segmentation (S4) can learn rich visual knowledge from low-cost unlabeled images. However, traditional S4 architectures all face the challenge of low-quality pseudo-labels, especially for the teacher-student framework.We propose a novel SemiEarth model that introduces vision-language models (VLMs) to address the S4 issues for the remote sensing (RS) domain. Specifically, we invent a VLM pseudo-label purifying (VLM-PP) structure to purify the teacher network's pseudo-labels, achieving substantial improvements. Especially in multi-class boundary regions of RS images, the VLM-PP module can significantly improve the quality of pseudo-labels generated by the teacher, thereby correctly guiding the student model's learning. Moreover, since VLM-PP equips VLMs with open-world capabilities and is independent of the S4 architecture, it can correct mispredicted categories in low-confidence pseudo-labels whenever a discrepancy arises between its prediction and the pseudo-label. We conducted extensive experiments on multiple RS datasets, which demonstrate that our SemiEarth achieves SOTA performance. More importantly, unlike previous SOTA RS S4 methods, our model not only achieves excellent performance but also offers good interpretability. The code is released at https://github.com/wangshanwen001/SemiEarth.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。