无需人工标注标签,用视觉一致性自动评估地理图像推理结果。
RemoteZero: Geospatial Reasoning with Zero Labels

- 用图像裁片与查询语义一致性做内在奖励,替代人工标签。
- 在无标签数据上表现超越监督基线,且随数据量增长持续优化。
- 适合大规模遥感数据的自动化地理推理,尤其适合缺乏标注场景。
地理空间推理需要模型识别满足复杂且常隐含用户意图的图像区域。现有强化学习方法虽无需手动标注推理轨迹,但仍需人类提供目标标签以构建奖励,限制了其在大规模未标注地球观测数据中的应用。本文提出 RemoteZero,一种基于强化学习的无标签地理空间推理框架。核心观察为:多模态大模型在评估候选区域时比生成解更可靠,而航拍图像能降低区域级验证的干扰。因此,RemoteZero 将每个预测区域转换为视觉裁片,并以该裁片与查询的语义一致性作为 GRPO 优化的内在奖励。该设计无需人工提供解标签,还支持通过复用前一轮模型作为验证器实现迭代自演化。实验表明,RemoteZero 在性能上优于强监督基线,可有效拓展至其他地球观测任务,且随着训练数据增加持续提升。我们希望这一方向能广泛惠及地球观测领域。
原文摘要 · Abstract (English)
Geospatial reasoning requires models to identify image regions that satisfy complex and often implicit user intents. Recent reinforcement learning approaches improve reasoning without manually annotated reasoning traces, but still require human-provided target labels to construct rewards, limiting their use on large-scale unlabeled Earth observation data. We introduce RemoteZero, a label-free framework for reinforcement-based geospatial reasoning. Our key observation is twofold: MLLMs are often more reliable at evaluating candidates than generating solutions, while aerial imagery reduces interference in region-level verification. Therefore, RemoteZero converts each predicted region into a visual crop and uses its semantic consistency with the query as an intrinsic reward for GRPO optimization. This formulation removes the need for human-provided solution labels and further supports iterative self-evolution by reusing previous-round models as verifiers. Experiments show that RemoteZero outperforms strong supervised baselines, applies effectively to other Earth observation tasks, and continues to improve through self-evolution as the training data expand. We hope this direction can broadly benefit the Earth observation community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。