arXiv:2512.17319cs.CVcs.AI2025-12被引 14

构建超高清遥感多模态模型评测基准,解决现有数据分辨率不足与任务设计缺陷问题。

A Benchmark for Ultra-High-Resolution Remote Sensing MLLMs

  • 设计覆盖九类感知与四类推理的多任务评估体系,支持多轮对话与多图交互。
  • 包含5329张超高清遥感图像,单图像素达3亿级,长边不少于4000像素。
  • 通过对抗过滤与人工验证,降低语言先验干扰,提升评测可信度。

多模态大语言模型在现有遥感基准上表现出强大的感知与推理能力。然而,多数先前基准依赖低分辨率图像,部分高分辨率基准存在推理任务设计缺陷。我们发现,仅依赖文本的LLM在无图像条件下仍可与多模态模型竞争,揭示当前基准与视觉理解评估目标间的严重不匹配。为此,我们提出RSHR-Bench,一个面向遥感视觉理解与推理的超高清基准。该基准包含5,329张全场景图像,单图长边不低于4,000像素,总像素达约3×10⁸,数据源自广泛使用的遥感语料库与无人机采集。设计四类任务:多项选择型VQA、开放式VQA、图像描述生成、单图评估。涵盖九类感知类别与四类推理类型,支持多轮对话与多图交互。为减少语言先验影响,采用强语言模型进行对抗过滤,并经严格人工验证。最终构建3,864个VQA任务、3,913个图像描述任务及500对完全人工撰写或验证的单图评估问答对。对开源、闭源及遥感专用多模态模型的评估显示,在超高清场景下性能差距依然显著。代码已开源。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) demonstrate strong perception and reasoning performance on existing remote sensing (RS) benchmarks. However, most prior benchmarks rely on low-resolution imagery, and some high-resolution benchmarks suffer from flawed reasoning-task designs. We show that text-only LLMs can perform competitively with multimodal vision-language models on RS reasoning tasks without access to images, revealing a critical mismatch between current benchmarks and the intended evaluation of visual understanding. To enable faithful assessment, we introduce RSHR-Bench, a super-high-resolution benchmark for RS visual understanding and reasoning. RSHR-Bench contains 5,329 full-scene images with a long side of at least 4,000 pixels, with up to about 3 x 10^8 pixels per image, sourced from widely used RS corpora and UAV collections. We design four task families: multiple-choice VQA, open-ended VQA, image captioning, and single-image evaluation. These tasks cover nine perception categories and four reasoning types, supporting multi-turn and multi-image dialog. To reduce reliance on language priors, we apply adversarial filtering with strong LLMs followed by rigorous human verification. Overall, we construct 3,864 VQA tasks, 3,913 image captioning tasks, and 500 fully human-written or verified single-image evaluation VQA pairs. Evaluations across open-source, closed-source, and RS-specific VLMs reveal persistent performance gaps in super-high-resolution scenarios. Code: https://github.com/Yunkaidang/RSHR

遥感多模态评测基准超高清

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。