arXiv:2505.15401cs.CV2025-05CVPR被引 9

多模态遥感图像让机器理解复杂问题,准确率达65.56%

Visual Question Answering on Multiple Remote Sensing Image Modalities

  • 融合高分辨率RGB、多光谱和雷达三种遥感图像
  • 在新数据集上实现65.56%的问答准确率
  • 适合遥感、医疗影像等多模态视觉研究者

视觉问答(VQA)中视觉特征提取至关重要。在遥感等领域,融合携带互补光谱、空间和上下文信息的不同图像模态可显著提升视觉表征能力。本文提出在遥感场景下引入多模态图像进行VQA,开创了新的研究任务。为此构建了TAMMI数据集,包含三类模态(超高分辨率RGB、多光谱影像、合成孔径雷达),并提供自动化扩展管道。同时提出基于VisualBERT的MM-RSVQA模型,通过可训练融合机制整合多模态图像与文本。初步实验表明,该方法在挑战性任务上达到65.56%的准确率。此项工作为计算机视觉社区开辟了多模态多分辨率遥感VQA新方向,并可推广至医学影像等领域。代码与数据集见https://tammi.sylvainlobry.com/。

原文摘要 · Abstract (English)

The extraction of visual features is an essential step in Visual Question Answering (VQA). Building a good visual representation of the analyzed scene is indeed one of the essential keys for the system to be able to correctly understand the latter in order to answer complex questions. In many fields such as remote sensing, the visual feature extraction step could benefit significantly from leveraging different image modalities carrying complementary spectral, spatial and contextual information. In this work, we propose to add multiple image modalities to VQA in the particular context of remote sensing, leading to a novel task for the computer vision community. To this end, we introduce a new VQA dataset, named TAMMI (Text and Multi-Modal Imagery) with diverse questions on scenes described by three different modalities (very high resolution RGB, multi-spectral imaging data and synthetic aperture radar). Thanks to an automated pipeline, this dataset can be easily extended according to experimental needs. We also propose the MM-RSVQA (Multi-modal Multi-resolution Remote Sensing Visual Question Answering) model, based on VisualBERT, a vision-language transformer, to effectively combine the multiple image modalities and text through a trainable fusion process. A preliminary experimental study shows promising results of our methodology on this challenging dataset, with an accuracy of 65.56% on the targeted VQA task. This pioneering work paves the way for the community to a new multi-modal multi-resolution VQA task that can be applied in other imaging domains (such as medical imaging) where multi-modality can enrich the visual representation of a scene. The dataset and code are available at https://tammi.sylvainlobry.com/.

遥感图像多模态视觉问答图像融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。