用中间步骤验证提升模型几何推理能力
Milestones over Outcome: Unlocking Geometric Reasoning with Sub-Goal Verifiable Reward
- 通过可验证的子目标评估替代结果导向监督
- 几何推理准确率提升9.7%,泛化至数学等任务
- 适合需要严谨逻辑推理的研究者
多模态大语言模型在复杂几何推理上表现不佳,主要因为基于最终结果的“黑箱”监督无法区分偶然猜对与严格推导。为此,我们提出子目标层面的评估与学习范式。首先构建GeoGoal基准,通过形式化验证数据引擎将抽象证明转化为可验证的数值子目标,揭示了推理质量与结果准确性之间的显著差异。基于此,我们提出子目标可验证奖励(SGVR)框架,以骨架率(Skeleton Rate)为基础,用密集奖励替代稀疏信号。实验表明,SGVR不仅使几何性能提升9.7%,还展现出强泛化能力,在通用数学任务上提升8.0%,其他通用推理任务提升2.8%,证明其在多个领域具有广泛适用性。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) struggle with complex geometric reasoning, largely because "black box" outcome-based supervision fails to distinguish between lucky guesses and rigorous deduction. To address this, we introduce a paradigm shift towards subgoal-level evaluation and learning. We first construct GeoGoal, a benchmark synthesized via a rigorous formal verification data engine, which converts abstract proofs into verifiable numeric subgoals. This structure reveals a critical divergence between reasoning quality and outcome accuracy. Leveraging this, we propose the Sub-Goal Verifiable Reward (SGVR) framework, which replaces sparse signals with dense rewards based on the Skeleton Rate. Experiments demonstrate that SGVR not only enhances geometric performance (+9.7%) but also exhibits strong generalization, transferring gains to general math (+8.0%) and other general reasoning tasks (+2.8%), demonstrating broad applicability across diverse domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。