构建超大分辨率遥感图像评测基准,检验多模态大模型感知与推理能力。
XLRS-Bench: Could Your Multimodal LLMs Understand Extremely Large Ultra-High-Resolution Remote Sensing Imagery?
- 设计超大尺寸遥感图像评测集,平均分辨率8500×8500像素。
- 人工精标注+半自动工具,提升标注质量与效率。
- 涵盖16项子任务,评估感知与时空推理等关键能力,适合遥感应用研究者。
多模态大语言模型(MLLMs)的迅猛发展亟需新基准来量化评估其能力、揭示局限性并指引未来方向。然而,遥感领域因图像具有极高的分辨率和复杂的语义关系,现有评测基准普遍存在图像尺寸过小、标注质量有限、评估维度不足等问题。为此,我们提出XLRS-Bench:一个面向超大分辨率遥感场景的综合性评测基准,其平均图像尺寸达8500×8500,为迄今最大。所有样本均经人工精细标注,并借助新型半自动标题生成器辅助完成。基于该基准,定义了16项子任务,评估MLLMs的10类感知能力和6类推理能力,重点考察支持真实决策的高级认知过程及时空变化捕捉能力。实验表明,通用与遥感专用模型在该基准上表现仍有显著提升空间。我们已开源XLRS-Bench,以推动更强大遥感大模型的发展。
原文摘要 · Abstract (English)
The astonishing breakthrough of multimodal large language models (MLLMs) has necessitated new benchmarks to quantitatively assess their capabilities, reveal their limitations, and indicate future research directions. However, this is challenging in the context of remote sensing (RS), since the imagery features ultra-high resolution that incorporates extremely complex semantic relationships. Existing benchmarks usually adopt notably smaller image sizes than real-world RS scenarios, suffer from limited annotation quality, and consider insufficient dimensions of evaluation. To address these issues, we present XLRS-Bench: a comprehensive benchmark for evaluating the perception and reasoning capabilities of MLLMs in ultra-high-resolution RS scenarios. XLRS-Bench boasts the largest average image size (8500$\times$8500) observed thus far, with all evaluation samples meticulously annotated manually, assisted by a novel semi-automatic captioner on ultra-high-resolution RS images. On top of the XLRS-Bench, 16 sub-tasks are defined to evaluate MLLMs' 10 kinds of perceptual capabilities and 6 kinds of reasoning capabilities, with a primary emphasis on advanced cognitive processes that facilitate real-world decision-making and the capture of spatiotemporal changes. The results of both general and RS-focused MLLMs on XLRS-Bench indicate that further efforts are needed for real-world RS applications. We have open-sourced XLRS-Bench to support further research in developing more powerful MLLMs for remote sensing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。