arXiv:2511.22645cs.CV2025-11被引 14

无需人工思维链标注,模型自动生成地理场景推理。

GeoZero: Incentivizing Reasoning from Scratch on Geospatial Scenes

  • 用自生成数据训练模型,避免人工标注思维链
  • 在多个遥感基准上超越现有最佳方法
  • 适合研究地理智能与零样本推理的学者

多模态大语言模型在地理场景理解方面发展迅速。近期研究通常通过精心构建的思维链(CoT)数据进行冷启动训练以提升遥感多模态模型的推理能力,但这不仅带来高昂的标注成本,还引入人类偏见,限制模型推理多样性。为此,我们提出GeoZero框架,使多模态大语言模型在无预设思维链监督下实现地理空间推理。具体地,我们构建了两个数据集:GeoZero-Instruct用于监督微调获取初步地理知识,GeoZero-Hard用于后续强化学习阶段激发深度推理。此外,我们提出答案锚定组相对策略优化(A²GRPO),通过模型自身答案对推理过程进行正则化,鼓励多样化且准确的思考。在多个遥感视觉-语言基准上的实验表明,GeoZero不仅超越现有最先进方法,还在多样地理任务中展现出普遍涌现的推理能力。代码、数据与模型已公开于https://github.com/MiliLab/GeoZero。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) have undergone rapid development in advancing geospatial scene understanding. Recent studies have sought to enhance the reasoning capabilities of remote sensing MLLMs, typically through cold-start training with elaborately curated chain-of-thought (CoT) data. However, this approach not only incurs substantial annotation costs but also introduces human biases that may limit the diversity of model reasoning. To address these challenges, we propose GeoZero, a framework that enables MLLMs to perform geospatial reasoning without any predefined CoT supervision. Specifically, we construct two datasets, GeoZero-Instruct and GeoZero-Hard. GeoZero-Instruct allows the model to acquire preliminary geospatial knowledge through supervised fine-tuning, while GeoZero-Hard stimulates deep reasoning during the subsequent reinforcement learning stage. Furthermore, we introduce Answer-Anchored Group Relative Policy Optimization (A$^2$GRPO), where the reasoning process is regularized by the model's own answers, encouraging diverse yet accurate thinking. Extensive experiments on multiple remote sensing vision-language benchmarks demonstrate that GeoZero not only surpasses existing state-of-the-art methods but also fosters universal emergent reasoning capabilities across diverse geospatial tasks. Code, data, and models are available at https://github.com/MiliLab/GeoZero.

地理推理零样本强化学习多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。