让模型像人一样有计划地看高分辨率遥感图,找证据不漏不重。
GeoVista: Visually Grounded Active Perception for Vision-Language Understanding of Ultra-High-Resolution Remote Sensing Images

- 先规划全局路径,再分枝检查候选区域,避免盲目搜索。
- 在三个遥感基准上达到当前最好性能,显著提升定位与问答准确率。
- 适合需要精细分析大场景图像的研究者,如城市规划、灾害监测。
解析超高清(UHR)遥感图像需在大范围场景中寻找稀疏且微小的视觉线索。现有遥感视觉语言模型虽能通过缩放和裁剪工具查看局部区域,但多数探索策略仅采用单次聚焦或单一顺序路径,易丢失全局上下文、遗漏分散区域,或重复访问相同证据。为此,我们提出GeoVista,一种面向超高清遥感理解的规划驱动主动感知框架。不同于固定缩放路径,GeoVista先构建全局探索计划,再通过分支式局部检查验证多个候选区域,并显式维护证据状态以实现跨区域聚合与去重。为支持此行为,我们引入APE-GRO——一个冷启动监督轨迹语料库,将多样化的UHR任务重构为统一、尺度不变的空间交互推理过程。我们还设计了观察-规划-追踪机制,实现全局观测、自适应区域检查与证据追踪,并采用基于GRPO的策略,通过分步奖励对齐模型在规划、定位与最终答案正确性上的表现。在RSHR-Bench、XLRS-Bench和LRS-VQA上的实验表明,GeoVista达到领先性能。代码与数据集见https://github.com/ryan6073/GeoVista。
原文摘要 · Abstract (English)
Interpreting ultra-high-resolution (UHR) remote sensing images requires models to search for sparse and tiny visual evidence across large-scale scenes. Existing remote sensing vision-language models can inspect local regions with zooming and cropping tools, but most exploration strategies follow either a one-shot focus or a single sequential trajectory. Such single-path exploration can lose global context, leave scattered regions unvisited, and revisit or count the same evidence multiple times. To this end, we propose GeoVista, a planning-driven active perception framework for UHR remote sensing interpretation. Instead of committing to one zooming path, GeoVista first builds a global exploration plan, then verifies multiple candidate regions through branch-wise local inspection, while maintaining an explicit evidence state for cross-region aggregation and de-duplication. To enable this behavior, we introduce APE-GRO, a cold-start supervised trajectory corpus that reformulates diverse UHR tasks as Global-Region-Object interactive reasoning processes with a unified, scale-invariant spatial representation. We further design an Observe-Plan-Track mechanism for global observation, adaptive region inspection, and evidence tracking, and align the model with a GRPO-based strategy using step-wise rewards for planning, localization, and final answer correctness. Experiments on RSHR-Bench, XLRS-Bench, and LRS-VQA show that GeoVista achieves state-of-the-art performance. Code and dataset are available at https://github.com/ryan6073/GeoVista.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。