arXiv:2602.14201cs.CVcs.AI2026-02被引 6

让大模型学会按需聚焦,精准识别高分辨率遥感图像中的关键信息。

GeoEyes: On-Demand Visual Focusing for Evidence-Grounded Understanding of Ultra-High-Resolution Remote Sensing Imagery

  • 构建分阶段训练框架,引导模型学会主动选择性放大目标区域。
  • 在超高清遥感问答任务中达到54.23%准确率,显著优于现有方法。
  • 适合需要精准视觉分析的遥感、地理信息等领域的研究者使用。

“以图思考”范式使多模态大语言模型能通过缩放工具主动探索视觉场景,这对任务线索稀疏且微小的超高清遥感视觉问答至关重要。然而我们发现现有缩放增强型多模态大模型存在一致失效模式:工具使用同质化,即缩放行为趋于无关任务的固定模式,限制了有效证据获取。为此,我们提出GeoEyes,一个包含两阶段的训练框架:(1) 冷启动监督微调数据集UHR Chain-of-Zoom(UHR-CoZ),覆盖多样化的缩放策略;(2) 基于智能体的强化学习方法AdaZoom-GRPO,显式奖励证据增益与答案改进。所获模型具备按需缩放与适时停止能力,在超高清遥感基准测试中表现优异,于XLRS-Bench上取得54.23%准确率。

原文摘要 · Abstract (English)

The "thinking-with-images" paradigm enables multimodal large language models (MLLMs) to actively explore visual scenes via zoom-in tools. This is essential for ultra-high-resolution (UHR) remote sensing VQA, where task-relevant cues are sparse and tiny. However, we observe a consistent failure mode in existing zoom-enabled MLLMs: Tool Usage Homogenization, where tool calls collapse into task-agnostic patterns, limiting effective evidence acquisition. To address this, we propose GeoEyes, a staged training framework consisting of (1) a cold-start SFT dataset, UHR Chain-of-Zoom (UHR-CoZ), which covers diverse zooming regimes, and (2) an agentic reinforcement learning method, AdaZoom-GRPO, that explicitly rewards evidence gain and answer improvement during zoom interactions. The resulting model learns on-demand zooming with proper stopping behavior and achieves substantial improvements on UHR remote sensing benchmarks, with 54.23% accuracy on XLRS-Bench.

遥感分析多模态模型视觉聚焦

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。