arXiv:2501.08114cs.CV2025-01被引 8

单阶段Transformer模型提升遥感变化描述精度与效率

Change Captioning in Remote Sensing: Evolution to SAT-Cap -- A Single-Stage Transformer Approach

  • 采用单阶段融合机制,降低计算复杂度
  • 在LEVIR-CC上达140.23%的CIDEr分数,优于现有方法
  • 适合需要高效精准遥感变化描述的科研与应用者

变化描述对准确解读多时相遥感数据至关重要,能以自然语言直观监测地球动态。现有方法面临两大挑战:多阶段融合导致计算开销大,且单幅图像语义提取不足影响物体描述细节。为此,本文提出基于Transformer的单阶段模型SAT-Cap,包含空间-通道注意力编码器、差异引导融合模块和描述解码器。相比传统需多阶段融合的架构,SAT-Cap仅用基于余弦相似度的简单融合模块,显著降低模型复杂度。通过联合建模空间与通道信息,显著增强对多时相遥感图像中目标的语义提取能力。大量实验验证其有效性,在LEVIR-CC数据集上获得140.23%的CIDEr得分,在DUBAI-CC上达97.74%,超越当前最优方法。代码与预训练模型将公开。

原文摘要 · Abstract (English)

Change captioning has become essential for accurately describing changes in multi-temporal remote sensing data, providing an intuitive way to monitor Earth's dynamics through natural language. However, existing change captioning methods face two key challenges: high computational demands due to multistage fusion strategy, and insufficient detail in object descriptions due to limited semantic extraction from individual images. To solve these challenges, we propose SAT-Cap based on the transformers model with a single-stage feature fusion for remote sensing change captioning. In particular, SAT-Cap integrates a Spatial-Channel Attention Encoder, a Difference-Guided Fusion module, and a Caption Decoder. Compared to typical models that require multi-stage fusion in transformer encoder and fusion module, SAT-Cap uses only a simple cosine similarity-based fusion module for information integration, reducing the complexity of the model architecture. By jointly modeling spatial and channel information in Spatial-Channel Attention Encoder, our approach significantly enhances the model's ability to extract semantic information from objects in multi-temporal remote sensing images. Extensive experiments validate the effectiveness of SAT-Cap, achieving CIDEr scores of 140.23% on the LEVIR-CC dataset and 97.74% on the DUBAI-CC dataset, surpassing current state-of-the-art methods. The code and pre-trained models will be available online.

遥感变化文本生成Transformer单阶段

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。