arXiv:2603.09551cs.CV2026-03被引 1

让遥感模型像解题一样一步步推理,还能自我验证。

GeoSolver: Scaling Test-Time Reasoning in Remote Sensing with Fine-Grained Process Supervision

  • 用细粒度过程监督训练奖励模型,确保每步推理都对得起图像。
  • 在多个遥感数据集上表现超越现有方法,90亿参数模型达顶尖水平。
  • 可通用提升其他视觉语言模型性能,适合做遥感智能分析的开发者。

尽管视觉语言模型(VLM)显著推进了遥感图像理解,但实现复杂、分步推理仍面临挑战。现有链式思维(CoT)方法虽有进展,但中间步骤的视觉一致性难以保证。为此,我们提出GeoSolver框架,将遥感推理转向可验证的过程监督强化学习。首先,通过熵引导的蒙特卡洛树搜索(MCTS)与定向视觉幻觉注入,构建大规模、标记级过程监督数据集Geo-PRM-2M。基于该数据集,训练出一个标记级过程奖励模型(GeoPRM),提供精细的视觉忠实度反馈。为有效利用这些验证信号,提出过程感知树状GRPO算法,结合树结构探索与忠实度加权奖励机制,精准分配中间步骤的信用。大量实验表明,所获模型GeoSolver-9B在多种遥感基准测试中达到最优表现。关键的是,GeoPRM实现了强大的测试时扩展(TTS)。作为通用地理空间验证器,它能无缝提升GeoSolver-9B性能,并直接增强通用VLM,展现卓越跨模型泛化能力。

原文摘要 · Abstract (English)

While Vision-Language Models (VLMs) have significantly advanced remote sensing interpretation, enabling them to perform complex, step-by-step reasoning remains highly challenging. Recent efforts to introduce Chain-of-Thought (CoT) reasoning to this domain have shown promise, yet ensuring the visual faithfulness of these intermediate steps remains a critical bottleneck. To address this, we introduce GeoSolver, a novel framework that transitions remote sensing reasoning toward verifiable, process-supervised reinforcement learning. We first construct Geo-PRM-2M, a large-scale, token-level process supervision dataset synthesized via entropy-guided Monte Carlo Tree Search (MCTS) and targeted visual hallucination injection. Building upon this dataset, we train GeoPRM, a token-level process reward model (PRM) that provides granular faithfulness feedback. To effectively leverage these verification signals, we propose Process-Aware Tree-GRPO, a reinforcement learning algorithm that integrates tree-structured exploration with a faithfulness-weighted reward mechanism to precisely assign credit to intermediate steps. Extensive experiments demonstrate that our resulting model, GeoSolver-9B, achieves state-of-the-art performance across diverse remote sensing benchmarks. Crucially, GeoPRM unlocks robust Test-Time Scaling (TTS). Serving as a universal geospatial verifier, it seamlessly scales the performance of GeoSolver-9B and directly enhances general-purpose VLMs, highlighting its remarkable cross-model generalization.

遥感推理强化学习验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。