用视觉语言模型提升遥感变化检测的语义理解与定位精度
ViLaCD-R1: A Vision-Language Framework for Semantic Change Detection in Remote Sensing
- 分两阶段:先用视觉语言模型生成粗略变化掩码,再融合特征精修边界
- 在多个基准上实现最优准确率,对非语义干扰有强鲁棒性
- 适合需要高精度变化定位的遥感应用,如城市变迁监测
遥感变化检测(RSCD)是一项复杂的多图像推理任务,传统方法依赖像素级算子或编码器-解码器网络,难以捕捉高层语义且易受非语义扰动影响。尽管近期基于多模态和视觉语言模型(VLM)的方法通过引入文本描述增强了变化区域的语义理解,但仍面临空间定位不准、像素级边界模糊及可解释性差等问题。为此,我们提出ViLaCD-R1,一个两阶段框架,包含多图像推理器(MIR)和掩码引导解码器(MGD)。具体而言,VLM通过监督微调(SFT)和强化学习(RL)在块级双时相推理任务上训练,输入双时相图像块,输出粗略变化掩码;随后,解码器融合双时相图像特征与该粗掩码,预测精确二值变化图。在多个RSCD基准上的全面评估表明,ViLaCD-R1显著提升真实语义变化的识别与定位能力,稳健抑制非语义变化,并在复杂真实场景中达到最先进准确率。
原文摘要 · Abstract (English)
Remote sensing change detection (RSCD), a complex multi-image inference task, traditionally uses pixel-based operators or encoder-decoder networks that inadequately capture high-level semantics and are vulnerable to non-semantic perturbations. Although recent multimodal and vision-language model (VLM)-based approaches enhance semantic understanding of change regions by incorporating textual descriptions, they still suffer from challenges such as inaccurate spatial localization, imprecise pixel-level boundary delineation, and limited interpretability. To address these issues, we propose ViLaCD-R1, a two-stage framework comprising a Multi-Image Reasoner (MIR) and a Mask-Guided Decoder (MGD). Specifically, the VLM is trained through supervised fine-tuning (SFT) and reinforcement learning (RL) on block-level dual-temporal inference tasks, taking dual-temporal image patches as input and outputting a coarse change mask. Then, the decoder integrates dual-temporal image features with this coarse mask to predict a precise binary change map. Comprehensive evaluations on multiple RSCD benchmarks demonstrate that ViLaCD-R1 substantially improves true semantic change recognition and localization, robustly suppresses non-semantic variations, and achieves state-of-the-art accuracy in complex real-world scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。