arXiv:2604.24031cs.CV2026-04

融合视觉结构与语义信息,提升遥感图像描述生成准确率

JSSFF: A Joint Structural-Semantic Fusion Framework for Remote Sensing Image Captioning

论文配图:JSSFF: A Joint Structural-Semantic Fusion Framework for Remote Sensing Image Captioning
图 1 · 摘自论文原文
  • 通过原图与边缘感知图联合编码,增强特征表达与边界感知
  • 在多个基准数据集上优于现有模型,定量与定性指标均领先
  • 适合遥感图像理解、智能地图生成等应用领域研究者参考

编码器-解码器框架已成为主流。编码器从输入图像中提取视觉特征,解码器则基于序列到序列建模生成文本描述。现有模型多关注解码阶段,而忽略图像中有效信息的提取,这对物体及其关系的理解至关重要。遥感图像结构复杂,存在遮挡、重叠和边缘模糊等问题,导致目标难以完整识别。为此,本文提出一种边缘感知融合方法,将原始图像与其边缘感知版本共同输入编码器,以增强特征表示与边界敏感度。同时采用基于对比的束搜索(CBBS)生成描述,通过公平比较候选句,在量化指标与描述相关性之间取得平衡。实验结果表明,本模型在多个基准数据集上均优于现有基线模型,无论在定量还是定性评价上均有显著优势。

原文摘要 · Abstract (English)

The encoder-decoder framework has become widely popular nowadays. In this model, the encoder extracts informative visual features from an input image, and the decoder employs a sequence-to-sequence formulation to generate the corresponding textual description from these features. The existing models focus more on the decision part. However, extracting meaningful information from the image can help the decoder generate an accurate caption by providing information about the objects and their relationship. Remote sensing images are highly complex. One major challenge is detecting objects that extend beyond their visible boundaries due to occlusion, overlapping structures, and unclear edges. Hence, there is a need to design an approach that can effectively capture both high-level semantics and low-level spatial details for accurate caption generation. In this work, we have proposed an edge-aware fusion method by incorporating the original image and its edge-aware version into the encoder to enhance feature representation and boundary awareness. We used a comparison-based beam search (CBBS) to generate captions to achieve a balanced trade-off between quantitative metrics and qualitative caption relevance through fairness-based comparison of candidate captions. Experimental results demonstrate our model's superiority over several baseline models in quantitative and qualitative perspectives.

遥感图像图像描述边缘感知序列生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。