arXiv:2409.18467cs.LG2024-09被引 12

用文本图卷积提升遥感图像描述生成质量

A TextGCN-Based Decoding Approach for Improving Remote Sensing Image Captioning

  • 引入文本图卷积网络捕捉词间语义关系
  • 在三个数据集上多个指标领先现有方法
  • 适合遥感智能分析与自动化描述场景

遥感图像在风险管控、安全监测和气象分析等领域具有重要价值,但人工标注成本高且需专业知识。本文提出一种基于文本图卷积网络(TextGCN)的编码器-解码器框架,通过TextGCN提取句子与语料级词间语义关系,增强解码器理解能力;同时采用基于对比的束搜索策略,保障生成过程公平性。我们在三个公开数据集上,使用七项指标(BLEU-1 至 BLEU-4、METEOR、ROUGE-L、CIDEr)进行评估,结果表明该方法显著优于当前主流编码器-解码器模型。

原文摘要 · Abstract (English)

Remote sensing images are highly valued for their ability to address complex real-world issues such as risk management, security, and meteorology. However, manually captioning these images is challenging and requires specialized knowledge across various domains. This letter presents an approach for automatically describing (captioning) remote sensing images. We propose a novel encoder-decoder setup that deploys a Text Graph Convolutional Network (TextGCN) and multi-layer LSTMs. The embeddings generated by TextGCN enhance the decoder's understanding by capturing the semantic relationships among words at both the sentence and corpus levels. Furthermore, we advance our approach with a comparison-based beam search method to ensure fairness in the search strategy for generating the final caption. We present an extensive evaluation of our approach against various other state-of-the-art encoder-decoder frameworks. We evaluated our method across three datasets using seven metrics: BLEU-1 to BLEU-4, METEOR, ROUGE-L, and CIDEr. The results demonstrate that our approach significantly outperforms other state-of-the-art encoder-decoder methods.

遥感图像图像描述图卷积生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。