arXiv:2601.04118cs.CV2026-01被引 10

让遥感视觉语言模型的推理过程与答案一致,减少错误逻辑。

GeoReason: Aligning Thinking And Answering In Remote Sensing Vision-Language Models Via Logical Consistency Reinforcement Learning

  • 用几何构造和专家知识生成4000条推理轨迹,训练模型正确推理。
  • 通过一致性强化学习,使模型答案与推理链匹配,避免位置捷径。
  • 适合需要高可靠空间推理的遥感应用,如城市规划、灾害评估。

遥感视觉语言模型的发展正从感知识别转向高层推理,以提升复杂空间任务中的认知可靠性。然而现有模型常出现逻辑幻觉,即正确答案基于错误推理链或依赖位置捷径而非空间逻辑,导致决策不可靠。为此,我们提出GeoReason框架,旨在同步内部思考与最终判断。首先构建了包含4000条推理轨迹的GeoReason-Bench数据集,由几何原型和专家知识合成。采用两阶段训练策略:(1)监督知识初始化,赋予模型推理语法与领域知识;(2)一致性感知强化学习,引入新颖的逻辑一致性奖励,通过选项排列策略惩罚逻辑漂移,确保决策基于可验证的推理路径。实验表明,该框架显著提升遥感视觉语言模型的认知可靠性与可解释性,在多项指标上达到当前最优表现。

原文摘要 · Abstract (English)

The evolution of Remote Sensing Vision-Language Models(RS-VLMs) emphasizes the importance of transitioning from perception-centric recognition toward high-level deductive reasoning to enhance cognitive reliability in complex spatial tasks. However, current models often suffer from logical hallucinations, where correct answers are derived from flawed reasoning chains or rely on positional shortcuts rather than spatial logic. This decoupling undermines reliability in strategic spatial decision-making. To address this, we present GeoReason, a framework designed to synchronize internal thinking with final decisions. We first construct GeoReason-Bench, a logic-driven dataset containing 4,000 reasoning trajectories synthesized from geometric primitives and expert knowledge. We then formulate a two-stage training strategy: (1) Supervised Knowledge Initialization to equip the model with reasoning syntax and domain expertise, and (2) Consistency-Aware Reinforcement Learning to refine deductive reliability. This second stage integrates a novel Logical Consistency Reward, which penalizes logical drift via an option permutation strategy to anchor decisions in verifiable reasoning traces. Experimental results demonstrate that our framework significantly enhances the cognitive reliability and interpretability of RS-VLMs, achieving state-of-the-art performance compared to other advanced methods.

遥感推理视觉语言模型逻辑一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。