首个专用于遥感复杂推理的多模态基准,推动智能解译发展。
VLRS-Bench: A Vision-Language Reasoning Benchmark for Remote Sensing
- 构建涵盖认知、决策、预测三维度的遥感推理框架
- 含2000个问答对,平均问题130词,覆盖14项任务与8个时相
- 融合地理先验与专家知识,提升场景真实性和推理深度
多模态大模型(MLLM)在复杂推理方面取得进展,但现有遥感(RS)基准仍偏重感知任务,如目标识别和场景分类,制约了其在高阶推理任务中的应用。为此,我们提出首个专注于复杂遥感推理的视觉语言推理基准(VLRS-Bench)。该基准覆盖认知、决策与预测三大核心维度,包含2000个问答对,平均问题长度为130.19词,涵盖14个任务及最多八个时间阶段。通过融合遥感领域先验知识与专家经验的专用构建流程,确保地理空间真实性与推理复杂性。实验表明,当前主流先进MLLM在该基准上存在显著瓶颈,为推进遥感领域的多模态推理研究提供了关键洞见。项目代码已开源:https://github.com/MiliLab/VLRS-Bench。
原文摘要 · Abstract (English)
Recent advancements in Multimodal Large Language Models (MLLMs) have enabled complex reasoning. However, existing remote sensing (RS) benchmarks remain heavily biased toward perception tasks, such as object recognition and scene classification. This limitation hinders the development of MLLMs for cognitively demanding RS applications. To address this, we propose a Vision Language ReaSoning Benchmark (VLRS-Bench), which is the first benchmark exclusively dedicated to complex RS reasoning. Structured across the three core dimensions of Cognition, Decision, and Prediction, VLRS-Bench comprises 2,000 question-answer pairs with an average question length of 130.19 words, spanning 14 tasks and up to eight temporal phases. VLRS-Bench is constructed via a specialized pipeline that integrates RS-specific priors and expert knowledge to ensure geospatial realism and reasoning complexity. Experimental results reveal significant bottlenecks in existing state-of-the-art MLLMs, providing critical insights for advancing multimodal reasoning within the remote sensing community. The project repository is available at https://github.com/MiliLab/VLRS-Bench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。