用间接奖励让模型学会零样本地理推理,无需大量标注数据。
Unlocking Zero-Shot Geospatial Reasoning via Indirect Rewards
- 用地理位置等元数据做间接奖励,驱动大规模强化学习。
- 在25个以上任务上实现优异的零样本迁移,部分超越全监督模型。
- 适合研究稀有领域建模、少样本推理与自监督学习的学者。
在罕见领域(如地理空间)训练稳健的视觉-语言模型受限于监督信号稀缺。尽管地理影像数据丰富,但任务直接标注远少于常见领域。本文验证一个重要结论:源自看似无关元数据的间接可验证奖励,足以激发广泛下游任务(25+)上的复杂且泛化性强的地理空间推理能力。我们提出Geo-R1作为该范式的实证实例。不同于依赖有限任务特定标注(即直接奖励),Geo-R1利用基于跨视图对齐元数据(地理定位信息)的可扩展、可验证间接代理奖励,实现大规模强化学习。此类间接奖励成功引导模型在多样任务中发现并内化零样本地理空间推理能力,在分布外基准上表现卓越,甚至在某些基准上超越完全监督的专业模型。这些结果表明,优化间接可验证奖励可能为利用海量未标注数据档案解锁稀有领域中的通用推理能力提供可扩展路径。代码已公开:https://github.com/miniHuiHui/Geo-R1。
原文摘要 · Abstract (English)
Training robust reasoning vision-language models (VLMs) in rare domains (such as geospatial) is fundamentally constrained by supervision scarcity. While raw geospatial imagery is abundant, the amount of task-direct supervision falls far behind that of common domains. In this work, we validate an important conclusion: indirect verifiable rewards, derived from seemingly unrelated metadata, are sufficient to induce sophisticated and generalizable geospatial reasoning across a wide range of downstream tasks (25+). We present Geo-R1 as one empirical instantiation of this paradigm. Rather than relying on limited task-specific annotations (i.e., direct rewards), Geo-R1 utilizes scalable, verifiable indirect proxy rewards based on cross-view alignment with metadata (geolocation information) to drive reinforcement learning at scale. Such indirect rewards successfully motivate the model to discover and internalize zero-shot geospatial reasoning across diverse tasks, achieving extraordinary zero-shot transfer on out-of-distribution benchmarks and even surpassing fully supervised specialists on certain benchmarks. These findings indicate that optimizing for indirect verifiable rewards may provide a scalable pathway to unlock generalized reasoning capabilities in rare domains with massive unlabeled data archives. Our code is availavle at: https://github.com/miniHuiHui/Geo-R1.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。