用语言桥接地图与影像语义差异,提升变化检测精度
LaVIDE: Language-Prompted Satellite Change Detection via Map-Image Alignment
- 以语言为中介对齐地图语义与图像细节
- 多类别任务IoU提升18.4%,单类别提升5.2%
- 适合需快速更新地图的城建、灾评场景
基于地图参考与最新影像的遥感变化检测,在缺乏早期影像时可实现地表的及时观测。然而,高层地图类别与底层图像细节之间的语义鸿沟阻碍了同质特征提取,影响变化检测中时间关联的鲁棒性。针对这一问题,本文提出新颖框架LaVIDE(Language-Prompted Satellite Change Detection via Map-Image Alignment),利用语言作为中介,弥合地图与图像间的语义差距。具体而言,引入受限提示学习生成上下文感知的文本提示,使地图语义与图像内容对齐;并设计对象感知嵌入增强策略,将形状、边界等物体级属性融入地图表示。上述组件在统一的语言-视觉特征空间中实现稳健的跨模态对齐。在DynamicEarthNet、HRSCD、BANDON和SECOND四个基准上的大量实验表明,LaVIDE显著优于现有先进方法,多类别变化检测任务中交并比(IoU)提升18.4%,单类别任务提升5.2%。该框架不仅提升了地图-影像变化检测的准确性,还为最小人工干预下的快速地图更新提供了可行方案,有望在城市规划、灾害评估和生态保护等领域产生广泛影响。代码与数据集已公开于:https://github.com/ShuGuoJ/LAVIDE.git。
原文摘要 · Abstract (English)
Remote sensing change detection based on a map reference and an up-to-date image boosts timely observation of the Earth's surface when earlier images are lacking for comparison. However, the semantic gap between high-level map categories and low-level image details hinders the extraction of homogeneous features for robust temporal association in change detection. Unlike conventional approaches that either compare pixel-level visual similarity or propagate segmentation errors, \textcolor{black}{we propose a novel framework, \underline{La}nguage-\underline{VI}sion \underline{D}iscriminator for d\underline{E}tecting changes, LaVIDE}, which bridges the semantic gap between high-level map categories and low-level image details using language as an intermediary. Specifically, we introduce {\it restricted prompt learning} to generate context-aware textual prompts that align map semantics with image content, and an {\it object-aware embedding enhancement} strategy to integrate object-level attributes (e.g., shape, boundary) into map representations. These components enable robust cross-modal alignment within a unified language-vision feature space. Extensive experiments on four benchmarks, DynamicEarthNet, HRSCD, BANDON, and SECOND, demonstrate that LaVIDE outperforms state-of-the-art methods by significant margins, achieving $18.4\%$ and $5.2\%$ improvements in IoU on multi-class and single-class change detection tasks, respectively. Our framework not only advances the accuracy of map-image change detection but also provides a practical solution for rapid map updating with minimal human intervention, promising broad impacts in urban planning, disaster assessment, and ecological conservation. Code and datasets are available at: https://github.com/ShuGuoJ/LAVIDE.git.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。