arXiv:2605.07178cs.CV2026-05

从遥感图像标签中提取结构化文本,提升变化检测精度

Masks Can Talk: Extracting Structured Text Information from Single-Modal Images for Remote Sensing Change Detection

论文配图:Masks Can Talk: Extracting Structured Text Information from Single-Modal Images for Remote Sensing Change Detection
图 1 · 摘自论文原文
  • 从标注掩膜自动生成包含位置、对象、方式、数量的四元组文本
  • 在新构建的Gaza-Change-v2数据集上,F_scd达66.14%,优于依赖大模型的多模态方法
  • 零额外标注成本,提供精确无噪的多模态监督,适合遥感变化分析场景

遥感变化检测对城市监测、灾情评估和资源管理至关重要。然而,单模态深度学习方法常将视觉相似但语义无关的变化误判为真实变化。现有多模态方法虽引入文本辅助监督,但描述或过于粗略无结构,或由模型生成且含噪声。关键在于,所有方法忽视了一个事实:每个变化检测数据集的标准标注掩膜已隐含精细变化语义——包括变化位置、变化前后地物类型、变化方式及涉及对象数量。本文提出S2M框架,直接从变化标签中获取结构化文本特征,无需额外标注。具体而言,每个变化区域被自动转写为(位置、对象、方式、数量)四元组,并生成固定模板文本,提供精确、密集且无噪声的多模态监督。采用两阶段训练策略:先在遥感图像上微调以获得领域特定表示,再引入双向对比损失的多模态解码器,实现视觉特征与结构化文本嵌入的深层对齐。为验证方法,构建Gaza-Change-v2数据集,该数据集涵盖加沙地带多类变化。在该数据集上,S2M取得Sek 17.80%和F_scd 66.14%的性能,显著超越依赖大语言模型的多模态方法。本工作证明:掩膜确实能‘说话’,它们精准告知了变化的‘谁、何处、如何、多少’。

原文摘要 · Abstract (English)

Remote sensing change detection is pivotal for urban monitoring, disaster assessment, and environmental resource management. Yet, unimodal deep learning methods frequently confuse genuine semantic changes with visually similar but irrelevant variations. Recent multimodal approaches incorporate text as auxiliary supervision, but their descriptions are either semantically coarse and unstructured or model-generated and thus noisy. Critically, all of them overlook a simple fact: fine-grained change semantics are already implicitly encoded in the ground-truth mask labels that come standard with every change detection dataset. These masks know where the change happened, what the land-cover types were before and after, how the transition occurred, and how many objects were involved. In this paper, we propose S2M, a framework that obtains structured textual features directly from change labels at zero additional annotation cost. Specifically, each change region is automatically transcribed into a semantic quadruple (where, what, how, how many) and converted into several fixed-template text descriptions, providing precise, dense, and noise-free multimodal supervision. We adopts a two-stage training strategy to fine-tune on remote sensing imagery firstly for robust domain-specific representation, after which a multimodal decoder with a bi-directional contrastive loss is introduced to achieve deep alignment between visual features and structured textual embeddings. To validate our method, we construct Gaza-Change-v2, a new multi-class change detection (MCD) dataset about the Gaza Strip. On this MCD dataset, S2M achieves a Sek of 17.80\% and an F$_{\text{scd}}$ of 66.14\%, notably surpassing even multimodal methods that leverage large language models. Our work demonstrates that masks can indeed talk. They tell us exactly what, where, how, and how many changes have occurred.

变化检测遥感图像结构化文本多模态学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。