arXiv:2509.22221cs.CV2025-09被引 18

让遥感视觉模型像人一样一步步推理,结果可验证。

Towards Faithful Reasoning in Remote Sensing: A Perceptually-Grounded GeoSpatial Chain-of-Thought for Vision-Language Models

  • 用分步推理链构建可验证的分析流程
  • 在Geo-CoT380k数据集上训练,准确率显著提升
  • 适合需要可信分析的遥感决策场景

遥感领域的视觉语言模型在复杂分析任务中表现不佳,主要因端到端训练跳过关键推理步骤,导致输出不可验证。为此,我们提出感知基础的地理空间思维链(Geo-CoT)框架,将遥感分析建模为可验证的多步过程。通过两阶段对齐策略实现该分析流程:首先使用监督微调(SFT)建立基础认知结构,再利用群体奖励策略优化(GRPO)提升推理政策的事实正确性。由此产生的模型RSThinker不仅输出最终答案,还提供可追溯、可验证的分析轨迹。该能力在多项任务中显著优于现有最优模型。论文公开发布Geo-CoT380k数据集和RSThinker模型,推动地球观测从模糊感知迈向结构化、可验证推理。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) in remote sensing often fail at complex analytical tasks, a limitation stemming from their end-to-end training paradigm that bypasses crucial reasoning steps and leads to unverifiable outputs. To address this limitation, we introduce the Perceptually-Grounded Geospatial Chain-of-Thought (Geo-CoT), a framework that models remote sensing analysis as a verifiable, multi-step process. We instill this analytical process through a two-stage alignment strategy, leveraging Geo-CoT380k, the first large-scale dataset of structured Geo-CoT rationales. This strategy first employs supervised fine-tuning (SFT) to instill the foundational cognitive architecture, then leverages Group Reward Policy Optimization (GRPO) to refine the model's reasoning policy towards factual correctness. The resulting model, RSThinker, outputs both a final answer and its justifying, verifiable analytical trace. This capability yields dominant performance, significantly outperforming state-of-the-art models across a comprehensive range of tasks. The public release of our Geo-CoT380k dataset and RSThinker model upon publication serves as a concrete pathway from opaque perception towards structured, verifiable reasoning for Earth Observation.

遥感分析思维链可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。