arXiv:2605.20006cs.AI2026-05被引 1

GeoX通过自博弈与可验证奖励,让模型自主学习地理空间推理能力。

GeoX: Mastering Geospatial Reasoning Through Self-Play and Verifiable Rewards

论文配图:GeoX: Mastering Geospatial Reasoning Through Self-Play and Verifiable Rewards
图 1 · 摘自论文原文
  • 用可执行程序生成并求解空间问题,实现多模式推理。
  • 在无标注数据下提升基线模型平均5.5分,媲美百万级标注训练。
  • 适合对地理信息分析、空间逻辑推理感兴趣的开发者和研究者。

地理空间推理需在复杂场景的空间结构上解决图像关联问题,但高昂的标注成本限制了其发展。我们提出GeoX,一种基于自博弈的框架,通过可执行程序生成并获得可验证奖励,无需大规模人工标注数据。给定卫星或航拍图像,该框架采用单一多模态策略,将空间问题转化为可执行程序,并在三种推理模式——溯因、演绎、归纳——下对空间基元与图像理解工具进行推理。验证器执行每个程序以生成奖励信号,通过强化学习联合优化生成与求解两个角色。GeoX在不依赖人工标注的情况下,平均提升基础视觉语言模型5.5分,达到或超过依赖数百万条标注数据的传统基线。同时,我们发布了由自博弈积累的地理空间理解基准数据集。

原文摘要 · Abstract (English)

Geospatial reasoning requires solving image-grounded problems over the complex spatial structure of a scene. However, developing this capability is hindered by the cost of annotating a vast and combinatorial question space. We propose GeoX, a self-play framework that acquires spatial logic through executable programs that yield verifiable rewards, without relying on large-scale human-curated data Given a satellite or aerial image, our framework employs a single multimodal policy that proposes spatial problems as executable programs and solves them under three reasoning modes-abduction, deduction, and induction-over spatial primitives and an image understanding tool. A verifier executes each program to covert a reward signal that jointly optimizes the two roles via reinforcement learning. GeoX consistently improves its base VLMs by up to 5.5 points on average, matching or exceeding conventional baselines trained on millions of curated data. Along-side the proposed method, we release a benchmark for geospatial understanding accumulated through self-play.

地理推理自博弈多模态可验证奖励

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。