提出主动感知框架ZoomEarth,高效处理超分辨率遥感图像。
ZoomEarth: Active Perception for Ultra-High-Resolution Geospatial Vision-Language Tasks
- 引入主动感知机制,动态聚焦信息丰富区域。
- 在17类问题上达到领先性能,零样本下跨3个基准表现优异。
- 可集成至云去噪、分割等任务,接口简洁易用。
超高分辨率(UHR)遥感图像包含丰富细粒度信息,但有效处理仍具挑战。现有动态分辨率与标记剪枝方法受限于被动感知范式,在获取更细视觉输入时冗余增加。本文探索一种新主动感知范式,使模型可回溯信息丰富区域。首先构建LRS-GRO——一个面向UHR遥感主动感知的大规模基准数据集,涵盖全球、区域、物体三级共17类问题,通过半自动流程标注。基于LRS-GRO,提出ZoomEarth自适应裁剪-缩放框架,采用新型区域引导奖励实现细粒度指导。经监督微调(SFT)与组相对策略优化(GRPO)训练,ZoomEarth在LRS-GRO上达当前最优,并在零样本设置下于三个公开UHR遥感基准表现领先。此外,可通过简单工具接口无缝集成至云去除、去噪、分割、图像编辑等下游任务,展现强通用性与扩展性。
原文摘要 · Abstract (English)
Ultra-high-resolution (UHR) remote sensing (RS) images offer rich fine-grained information but also present challenges in effective processing. Existing dynamic resolution and token pruning methods are constrained by a passive perception paradigm, suffering from increased redundancy when obtaining finer visual inputs. In this work, we explore a new active perception paradigm that enables models to revisit information-rich regions. First, we present LRS-GRO, a large-scale benchmark dataset tailored for active perception in UHR RS processing, encompassing 17 question types across global, region, and object levels, annotated via a semi-automatic pipeline. Building on LRS-GRO, we propose ZoomEarth, an adaptive cropping-zooming framework with a novel Region-Guided reward that provides fine-grained guidance. Trained via supervised fine-tuning (SFT) and Group Relative Policy Optimization (GRPO), ZoomEarth achieves state-of-the-art performance on LRS-GRO and, in the zero-shot setting, on three public UHR remote sensing benchmarks. Furthermore, ZoomEarth can be seamlessly integrated with downstream models for tasks such as cloud removal, denoising, segmentation, and image editing through simple tool interfaces, demonstrating strong versatility and extensibility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。