arXiv:2506.09429cs.CV2025-06

轻量级Transformer提升遥感图像描述质量,兼顾边缘细节与部署效率

A Novel Lightweight Transformer with Edge-Aware Fusion for Remote Sensing Image Captioning

  • 压缩编码器维度并用GPT-2蒸馏版解码器降低计算开销
  • 在多个基准数据集上生成描述的BLEU-4和CIDEr得分均优于现有方法
  • 融合边缘感知模块,增强对遥感图像中物体边界等细粒度结构的理解

基于Transformer的模型在遥感图像描述任务中表现优异,能捕捉长距离依赖和上下文信息。但其高计算成本限制了实际部署,尤其在多模态框架中采用独立编码器与解码器时更为明显。现有模型多关注高层语义,常忽略边缘、轮廓和物体边界等细粒度结构特征。为此,本文提出一种轻量级Transformer架构:通过降低编码器层维度,并采用蒸馏版GPT-2作为解码器,结合知识蒸馏策略将复杂教师模型的知识迁移至轻量网络。此外,引入边缘感知增强策略,提升图像表征能力与物体边界理解,使模型更准确捕捉遥感图像中的细粒度空间细节。实验表明,该方法在多个基准数据集上显著提升描述质量,性能优于当前最优方法。

原文摘要 · Abstract (English)

Transformer-based models have achieved strong performance in remote sensing image captioning by capturing long-range dependencies and contextual information. However, their practical deployment is hindered by high computational costs, especially in multi-modal frameworks that employ separate transformer-based encoders and decoders. In addition, existing remote sensing image captioning models primarily focus on high-level semantic extraction while often overlooking fine-grained structural features such as edges, contours, and object boundaries. To address these challenges, a lightweight transformer architecture is proposed by reducing the dimensionality of the encoder layers and employing a distilled version of GPT-2 as the decoder. A knowledge distillation strategy is used to transfer knowledge from a more complex teacher model to improve the performance of the lightweight network. Furthermore, an edge-aware enhancement strategy is incorporated to enhance image representation and object boundary understanding, enabling the model to capture fine-grained spatial details in remote sensing images. Experimental results demonstrate that the proposed approach significantly improves caption quality compared to state-of-the-art methods.

遥感图像轻量模型边缘感知图像描述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。