arXiv:2504.04988cs.CVcs.AI2025-04中稿 · IEEE Geoscience an…被引 6

用卫星图像+外部知识,让遥感模型回答更精准的复杂问题。

Remote Sensing Retrieval-Augmented Generation: Bridging Remote Sensing Imagery and Comprehensive Knowledge with a Multi-Modal Dataset and Retrieval-Augmented Generation Model

  • 构建多模态遥感知识库,融合卫星图与全球地标文本。
  • 引入检索增强生成,根据图像或文字查询调用外部知识。
  • 在图像描述、分类、问答任务中显著优于现有方法。

近年来,视觉语言模型(VLM)在自然图像领域展现出强大能力。受此启发,遥感领域开始采用VLM进行场景理解、图像描述和视觉问答等任务。然而,现有遥感VLM多依赖封闭集场景理解,仅提供通用描述,缺乏外部知识接入能力,难以应对涉及领域专有或世界知识的复杂语义推理问题。为此,我们首次构建了包含14,141个全球知名地标的高分辨率卫星影像与详细文本描述的多模态遥感世界知识(RSWK)数据集,融合遥感领域知识与广义世界知识。基于该数据集,提出一种新型遥感检索增强生成(RS-RAG)框架,包含两个核心模块:多模态知识向量库构建模块,将遥感图像与关联文本编码至统一向量空间;知识检索与响应生成模块,根据图像和/或文本查询检索并重排序相关知识,将其融入知识增强提示,引导VLM生成语境相关的输出。我们在图像描述、图像分类和视觉问答三个代表性任务上验证了该方法的有效性,结果表明RS-RAG显著优于当前最优基线。

原文摘要 · Abstract (English)

Recent progress in VLMs has demonstrated impressive capabilities across a variety of tasks in the natural image domain. Motivated by these advancements, the remote sensing community has begun to adopt VLMs for remote sensing vision-language tasks, including scene understanding, image captioning, and visual question answering. However, existing remote sensing VLMs typically rely on closed-set scene understanding and focus on generic scene descriptions, yet lack the ability to incorporate external knowledge. This limitation hinders their capacity for semantic reasoning over complex or context-dependent queries that involve domain-specific or world knowledge. To address these challenges, we first introduced a multimodal Remote Sensing World Knowledge (RSWK) dataset, which comprises high-resolution satellite imagery and detailed textual descriptions for 14,141 well-known landmarks from 175 countries, integrating both remote sensing domain knowledge and broader world knowledge. Building upon this dataset, we proposed a novel Remote Sensing Retrieval-Augmented Generation (RS-RAG) framework, which consists of two key components. The Multi-Modal Knowledge Vector Database Construction module encodes remote sensing imagery and associated textual knowledge into a unified vector space. The Knowledge Retrieval and Response Generation module retrieves and re-ranks relevant knowledge based on image and/or text queries, and incorporates the retrieved content into a knowledge-augmented prompt to guide the VLM in producing contextually grounded responses. We validated the effectiveness of our approach on three representative vision-language tasks, including image captioning, image classification, and visual question answering, where RS-RAG significantly outperformed state-of-the-art baselines.

遥感检索增强多模态知识图谱

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。