对比12种CNN在遥感图像描述中的表现,找到最佳编码器提升生成质量。
Good Representation, Better Explanation: Role of Convolutional Neural Networks in Transformer-Based Remote Sensing Image Captioning
- 在Transformer框架中测试12种CNN作为视觉编码器
- 特定CNN使描述质量显著提升,人类评估验证效果
- 适合关注遥感图像生成与模型优化的研究者
遥感图像描述(RSIC)旨在从遥感图像生成有意义的文本描述。近年来,基于编码器-解码器的模型成为主流,其中编码器负责提取图像关键视觉特征并转化为紧凑表示,解码器则基于此生成连贯文本。尽管基于Transformer的解码器已广泛研究,编码器仍缺乏系统性探索。本文系统评估了12种不同卷积神经网络(CNN)架构在Transformer编码器框架中的表现,分为两个阶段:首先进行数值分析,按性能将CNN分类;最优候选者由人工标注员进行人类中心视角评估。此外,分析了贪婪搜索与束搜索对生成结果的影响。结果表明,编码器选择对生成质量具有决定性影响,特定CNN架构能显著提升遥感图像描述的准确性与丰富性。本研究为改进基于Transformer的图像描述模型提供了重要实证依据。
原文摘要 · Abstract (English)
Remote Sensing Image Captioning (RSIC) is the process of generating meaningful descriptions from remote sensing images. Recently, it has gained significant attention, with encoder-decoder models serving as the backbone for generating meaningful captions. The encoder extracts essential visual features from the input image, transforming them into a compact representation, while the decoder utilizes this representation to generate coherent textual descriptions. Recently, transformer-based models have gained significant popularity due to their ability to capture long-range dependencies and contextual information. The decoder has been well explored for text generation, whereas the encoder remains relatively unexplored. However, optimizing the encoder is crucial as it directly influences the richness of extracted features, which in turn affects the quality of generated captions. To address this gap, we systematically evaluate twelve different convolutional neural network (CNN) architectures within a transformer-based encoder framework to assess their effectiveness in RSIC. The evaluation consists of two stages: first, a numerical analysis categorizes CNNs into different clusters, based on their performance. The best performing CNNs are then subjected to human evaluation from a human-centric perspective by a human annotator. Additionally, we analyze the impact of different search strategies, namely greedy search and beam search, to ensure the best caption. The results highlight the critical role of encoder selection in improving captioning performance, demonstrating that specific CNN architectures significantly enhance the quality of generated descriptions for remote sensing images. By providing a detailed comparison of multiple encoders, this study offers valuable insights to guide advances in transformer-based image captioning models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。