构建首个面向罕见遥感图像的多任务评测基准,揭示现有模型短板。
RRS-10K: A Multitask Vision-Language Model Benchmark for Rare Remote Sensing Image Interpretation

- 构建涵盖军事场景的10,738张遥感图与多格式问答对
- 52个模型测试显示零样本性能仅中等,视觉定位能力弱
- 适合关注遥感智能分析、长尾场景研究的研究者
视觉语言模型(VLMs)在通用遥感任务上表现优异,但对罕见场景的理解能力仍不充分,因现有基准以常见城乡影像为主。为此,我们提出RRS-10K,一个面向罕见遥感图像解释的多任务基准。该数据集包含10,738张军事相关遥感图像及对应的多格式问答对,所有图像均来自一手来源,按感知、推理、鲁棒性三个维度划分为六个子维度和二十个细粒度任务。为提升选择题质量,构建过程中引入基于相似性的干扰项过滤策略(SDFS)。我们进一步评估了52个代表性模型,发现当前VLMs在罕见遥感图像零样本理解上表现仅中等,尤其在视觉定位、指代分割和复杂语义推理任务中存在明显缺陷。RRS-10K支持对长尾遥感理解失败模式的系统性分析,并为开发更可靠的遥感VLM提供指导。
原文摘要 · Abstract (English)
Vision-language models (VLMs) have achieved strong performance on general remote sensing tasks. However, their capability for rare scenes remains insufficiently understood, because existing benchmarks are dominated by common urban and rural imagery. To address this gap, we present RRS-10K, a benchmark for rare remote sensing image interpretation. RRS-10K contains 10,738 military-related remote sensing images and corresponding multiple format question-answer pairs for comprehensive evaluation. All of the images are collected from first-hand sources and organized into three capability dimensions, six sub-dimensions, and 20 leaf tasks, covering perception, reasoning, and robustness. To improve the quality of multiple-choice questions, we introduce a similarity-based distractor filtering strategy (SDFS) during benchmark construction. We further evaluate 52 representative models and show that current VLMs achieve only moderate zero-shot performance on rare remote sensing image interpretation, with clear weaknesses in visual grounding, referring segmentation, and complex semantic reasoning tasks. RRS-10K enables systematic analysis of failure modes in long-tail remote sensing interpretation and provides guidance for developing more reliable remote sensing VLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。