破解遥感大图中的微小目标误判难题,让模型看得更准。
UHR-Micro: Diagnosing and Mitigating the Resolution Illusion in Earth Observation VLMs

- 构建1.2万张高分辨率遥感图的诊断数据集,专测微小目标识别
- 发现现有模型在微小目标定位上错误率超60%,根源是缺乏引导
- 提出主动寻证框架,让模型像人一样逐步检查关键区域
视觉语言模型(VLM)日益应用于超高分辨率(UHR)遥感影像,但其在宏观场景与微观目标间存在严重尺度失配。我们称此现象为“分辨率幻觉”:更高分辨率看似提供更多细节,却未必提升对空间微小、任务相关证据的可靠感知。为此,我们提出UHR-Micro基准,包含11,253条基于1,212张UHR图像的指令,用于评估模型在遥感图像空间极限下的表现。该基准涵盖多样化的微目标尺度、上下文需求、任务类型和视觉条件,并提供诊断标注,支持可控评估与细粒度错误归因。实验表明,代表性高分辨率VLM在空间定位和证据解析上仍存在严重失败,即使拥有高分辨率输入。进一步分析显示,单纯增加模型容量无法解决该问题,根本原因在于缺乏对任务相关微证据的定位与使用引导。基于此,我们提出微观证据主动感知(MAP)框架,通过将查询分解为证据搜索步骤,主动检测候选区域,并基于局部观测进行回答。该方法使高分辨率推理从图像中心转向证据中心,显著提升微尺度感知能力。数据集与代码已开源。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) increasingly operate on ultra-high-resolution (UHR) Earth observation imagery, yet they remain vulnerable to a severe scale mismatch between large-scale scene context and micro-scale targets. We refer to this empirical gap as a "resolution illusion": higher input resolution provides the appearance of richer visual detail, but does not necessarily yield reliable perception of spatially small, task-relevant evidence. To benchmark this challenge, we introduce UHR-Micro, a benchmark comprising 11,253 instructions grounded in 1,212 UHR images, designed to evaluate VLMs at the spatial limits of native Earth observation imagery. UHR-Micro spans diverse micro-target scales, context requirements, task families, and visual conditions, and provides diagnostic annotations that support controlled evaluation and fine-grained error attribution. Experiments with representative high-resolution VLMs show substantial failures in spatial grounding and evidence parsing, despite access to high-resolution inputs. Further analysis suggests that these failures are not fully resolved by increasing model capacity, but are closely tied to insufficient guidance in locating and using task-relevant micro-evidence. Motivated by this finding, we propose Micro-evidence Active Perception (MAP), a reference agent that decomposes queries into evidence-seeking steps, actively inspects candidate regions, and grounds its answers in localized observations. MAP-Agent improves micro-level perception by making high-resolution reasoning evidence-centered rather than image-centered. Together, UHR-Micro and MAP-Agent provide a diagnostic platform for evaluating, understanding, and advancing high-resolution reasoning in Earth observation VLMs. Datasets and source code were released at https://github.com/MiliLab/UHR-Micro.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。