arXiv:2503.23453cs.CV2025-03被引 9

通过动态图优化提升遥感图像描述生成的语义与空间精度

Semantic-Spatial Feature Fusion with Dynamic Graph Refinement for Remote Sensing Image Captioning

  • 融合多层级视觉特征,结合文本语义增强图像理解
  • 动态图注意力机制优先关注场景相关物体,抑制无关信息
  • 在三个基准数据集上显著优于现有方法,适合遥感智能分析

遥感图像描述生成旨在生成与图像视觉特征紧密关联的语义准确描述。现有方法通常侧重于细粒度视觉特征提取和全局信息捕捉,但常忽略文本信息对视觉语义的补充作用,并难以精确定位与上下文最相关的物体。为此,本文提出语义-空间特征融合与动态图优化(SFDR)方法,集成语义-空间特征融合(SSFF)模块与动态图特征优化(DGFR)模块。SSFF模块利用预训练CLIP特征、网格特征和区域建议特征,实现多层次特征表示以融合丰富语义与空间信息。DGFR模块采用图注意力网络捕捉特征节点间关系,并通过动态加权机制突出当前场景中最相关的物体,抑制次要对象。实验结果表明,该方法在三个基准数据集上均取得显著性能提升,源代码将公开于https://github.com/zxk688。

原文摘要 · Abstract (English)

Remote sensing image captioning aims to generate semantically accurate descriptions that are closely linked to the visual features of remote sensing images. Existing approaches typically emphasize fine-grained extraction of visual features and capturing global information. However, they often overlook the complementary role of textual information in enhancing visual semantics and face challenges in precisely locating objects that are most relevant to the image context. To address these challenges, this paper presents a semantic-spatial feature fusion with dynamic graph refinement (SFDR) method, which integrates the semantic-spatial feature fusion (SSFF) and dynamic graph feature refinement (DGFR) modules. The SSFF module utilizes a multi-level feature representation strategy by leveraging pre-trained CLIP features, grid features, and ROI features to integrate rich semantic and spatial information. In the DGFR module, a graph attention network captures the relationships between feature nodes, while a dynamic weighting mechanism prioritizes objects that are most relevant to the current scene and suppresses less significant ones. Therefore, the proposed SFDR method significantly enhances the quality of the generated descriptions. Experimental results on three benchmark datasets demonstrate the effectiveness of the proposed method. The source code will be available at https://github.com/zxk688}{https://github.com/zxk688.

遥感图像图像描述动态图特征融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。