用视觉语言模型评估遥感图像语义分割质量,无需人工标注。
Remote Sensing Semantic Segmentation Quality Assessment based on Vision Language Model
- 基于预训练视觉语言模型,无监督提取分割质量特征。
- 在四个数据集上构建新数据集,8种方法结果用于标注质量。
- 相比现有模型显著提升评估准确率,适合遥感应用落地。
遥感影像语义分割在真实场景中受场景复杂性和图像质量差异影响,性能波动大,难以评估。现有指标依赖专家标注的物体级标签,不适用于此类场景。为此,我们提出基于视觉语言模型的无监督评估框架RS-SQA。该框架利用预训练的遥感专用视觉语言模型CLIP-RS进行语义理解,并通过分割模型的中间特征提取隐含质量信息。CLIP-RS通过净化文本训练,降低噪声,增强遥感领域语义表征能力;特征可视化显示其可有效区分不同分割质量水平。采用语义引导方式融合语义特征与低层分割特征,提升评估精度。为进一步支持研究,我们构建了RS-SQED数据集,从四个主流遥感分割数据集采样,并基于8种代表性分割方法的推理结果标注分割精度。实验表明,RS-SQA在该数据集上显著优于现有最先进模型,为预测分割准确率和高质量语义解释提供了有力支持,具有重要实用价值。
原文摘要 · Abstract (English)
The complexity of scenes and variations in image quality result in significant variability in the performance of semantic segmentation methods of remote sensing imagery (RSI) in supervised real-world scenarios. This makes the evaluation of semantic segmentation quality in such scenarios an issue to be resolved. However, most of the existing evaluation metrics are developed based on expert-labeled object-level annotations, which are not applicable in such scenarios. To address this issue, we propose RS-SQA, an unsupervised quality assessment model for RSI semantic segmentation based on vision language model (VLM). This framework leverages a pre-trained RS VLM for semantic understanding and utilizes intermediate features from segmentation methods to extract implicit information about segmentation quality. Specifically, we introduce CLIP-RS, a large-scale pre-trained VLM trained with purified text to reduce textual noise and capture robust semantic information in the RS domain. Feature visualizations confirm that CLIP-RS can effectively differentiate between various levels of segmentation quality. Semantic features and low-level segmentation features are effectively integrated through a semantic-guided approach to enhance evaluation accuracy. To further support the development of RS semantic segmentation quality assessment, we present RS-SQED, a dedicated dataset sampled from four major RS semantic segmentation datasets and annotated with segmentation accuracy derived from the inference results of 8 representative segmentation methods. Experimental results on the established dataset demonstrate that RS-SQA significantly outperforms state-of-the-art quality assessment models. This provides essential support for predicting segmentation accuracy and high-quality semantic segmentation interpretation, offering substantial practical value.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。