提出新型Transformer架构,提升遥感图像自动描述生成效果
SEMT: Static-Expansion-Mesh Transformer Network Architecture for Remote Sensing Image Captioning
- 融合静态扩展、记忆增强自注意力与网格Transformer结构
- 在UCM-Caption和NWPU-Caption数据集上超越现有最佳模型
- 适合遥感影像分析、环境监测等实际应用领域
图像字幕生成作为计算机视觉与自然语言处理交叉的重要任务,可实现从视觉内容自动生成描述性文本。在遥感领域,该技术对解析海量复杂的卫星影像具有重要意义,支持环境监测、灾害评估与城市规划等应用。本文提出一种基于Transformer的遥感图像字幕生成网络架构(SEMT),综合评估并整合了静态扩展、记忆增强自注意力与网格Transformer等技术。我们在UCM-Caption与NWPU-Caption两个基准遥感图像数据集上进行实验,所提最优模型在多数评价指标上优于当前最先进系统,展现出在真实遥感图像系统中应用的潜力。
原文摘要 · Abstract (English)
Image captioning has emerged as a crucial task in the intersection of computer vision and natural language processing, enabling automated generation of descriptive text from visual content. In the context of remote sensing, image captioning plays a significant role in interpreting vast and complex satellite imagery, aiding applications such as environmental monitoring, disaster assessment, and urban planning. This motivates us, in this paper, to present a transformer based network architecture for remote sensing image captioning (RSIC) in which multiple techniques of Static Expansion, Memory-Augmented Self-Attention, Mesh Transformer are evaluated and integrated. We evaluate our proposed models using two benchmark remote sensing image datasets of UCM-Caption and NWPU-Caption. Our best model outperforms the state-of-the-art systems on most of evaluation metrics, which demonstrates potential to apply for real-life remote sensing image systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。