构建首个面向遥感视觉语言模型的高质量问答数据集,提升模型理解与推理能力评估标准。
RSVLM-QA: A Benchmark Dataset for Remote Sensing Vision Language Model-based Question Answering
- 用大模型自动生成图文描述、空间关系和语义标签,结合真实分割数据计数
- 包含13,820张图像、16.2万组问答对,覆盖多样问题类型与复杂推理场景
- 适合遥感智能解译、视觉语言模型评测的研究者使用
遥感视觉问答(RS VQA)对地球观测数据的理解至关重要。现有数据集在标注丰富性、问题多样性及特定推理能力评估方面存在局限。本文提出RSVLM-QA,一个大规模、内容丰富的遥感领域视觉问答数据集。该数据集整合了WHU、LoveDA、INRIA和iSAID等主流遥感分割与检测数据集,采用双轨自动化标注流程:首先利用GPT-4.1结合精心设计提示词生成图像描述、空间关系、语义标签及基于描述的问答对;其次针对遥感图像中目标计数难题,从原始分割数据直接提取物体数量,再由GPT-4.1生成自然语言答案,并匹配预设问题模板形成计数类问答对。最终数据集包含13,820张图像和162,373组问答对,具备详尽标注与多样化问题类型。我们进行了数据集统计分析并与现有基准对比,凸显其标注深度与广度优势。同时在六种主流视觉语言模型上开展基准测试,验证了其对当前模型理解与推理能力的有效评估作用。我们认为RSVLM-QA将推动遥感视觉语言模型研究发展。
原文摘要 · Abstract (English)
Visual Question Answering (VQA) in remote sensing (RS) is pivotal for interpreting Earth observation data. However, existing RS VQA datasets are constrained by limitations in annotation richness, question diversity, and the assessment of specific reasoning capabilities. This paper introduces RSVLM-QA dataset, a new large-scale, content-rich VQA dataset for the RS domain. RSVLM-QA is constructed by integrating data from several prominent RS segmentation and detection datasets: WHU, LoveDA, INRIA, and iSAID. We employ an innovative dual-track annotation generation pipeline. Firstly, we leverage Large Language Models (LLMs), specifically GPT-4.1, with meticulously designed prompts to automatically generate a suite of detailed annotations including image captions, spatial relations, and semantic tags, alongside complex caption-based VQA pairs. Secondly, to address the challenging task of object counting in RS imagery, we have developed a specialized automated process that extracts object counts directly from the original segmentation data; GPT-4.1 then formulates natural language answers from these counts, which are paired with preset question templates to create counting QA pairs. RSVLM-QA comprises 13,820 images and 162,373 VQA pairs, featuring extensive annotations and diverse question types. We provide a detailed statistical analysis of the dataset and a comparison with existing RS VQA benchmarks, highlighting the superior depth and breadth of RSVLM-QA's annotations. Furthermore, we conduct benchmark experiments on Six mainstream Vision Language Models (VLMs), demonstrating that RSVLM-QA effectively evaluates and challenges the understanding and reasoning abilities of current VLMs in the RS domain. We believe RSVLM-QA will serve as a pivotal resource for the RS VQA and VLM research communities, poised to catalyze advancements in the field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。