arXiv:2607.25993cs.CV2026-07

让遥感大模型学会用多种工具分析超高清卫星图,比单纯放大更聪明。

Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing

论文配图:Beyond Zooming: Learning Multi-Tool Visual Reasoning for Ultra-High-Resolution Remote Sensing
图 1 · 摘自论文原文
  • 设计多工具协同推理框架,支持定位、比较、路径规划等复杂操作
  • 在1.3万张超高清遥感图上训练,准确率显著优于单工具放大方案
  • 适合需要全局分析的高阶遥感任务,如城市规划、灾害评估

超高清遥感影像为城市尺度的地球观测提供精细证据,但对多模态大语言模型构成挑战:相关证据常稀疏、局部且分散于极大视觉上下文中。现有方案依赖缩放工具进行局部检查,但在需要全局搜索、多区域对比、路径规划或分散证据推理的难题上效果饱和。为此,我们构建了包含1.3万组超高清遥感视觉问答样本的GeoMTVR数据集,涵盖交错的推理轨迹、多样化的视觉工具调用与返回观察结果,使模型可学习问题分解、工具选择、区域检查、对象定位、辅助推理及跨工具证据融合。结合监督微调与面向工具注意力的强化学习算法,我们提出GeoLens多工具视觉推理模型,显著优于直接推理和单工具缩放基线,在准确性、证据定位与工具使用效率上均有提升。

原文摘要 · Abstract (English)

Ultra-high-resolution (UHR) remote-sensing (RS) imagery provides fine-grained Earth-observation evidence over city-scale scenes, but poses a fundamental challenge for multimodal large language models (MLLMs): task-relevant evidence is often sparse, local, and spatially dispersed across extremely large visual contexts. A natural solution is to equip MLLMs with zoom-in tools for active local inspection. However, through a pilot study on XLRS-Bench, we find that zoom-in is only partially effective: it resolves easy and medium-level tasks with locally recoverable evidence, but saturates on hard cases requiring global search, multi-region comparison, path planning, or dispersed-evidence reasoning. Motivated by this finding, we move beyond single-tool zoom-in and introduce GeoMTVR, a large-scale Geospatial Multi-Tool Visual Reasoning dataset built from wide-area satellite imagery. GeoMTVR contains 13K UHR VQA samples with interleaved reasoning trajectories, diverse visual tool calls, and returned visual observations, enabling models to learn question decomposition, tool selection, regional inspection, object-level grounding, auxiliary visual reasoning, and cross-tool evidence integration. Beyond supervised fine-tuning, we propose a tool-attention-focused reinforcement learning algorithm that concentrates optimization on critical tool-use decisions, including when to invoke tools, which tool to select, where to apply it, and how to interpret tool outputs. By combining SFT on GeoMTVR with our RL algorithm, we develop GeoLens, a multi-tool visual reasoning MLLM for UHR RS. Experiments show that GeoLens consistently outperforms direct reasoning and single-tool zoom-in baselines, achieving stronger accuracy, better evidence grounding, and more efficient tool-use trajectories.

遥感分析多工具推理大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。