通过一致性验证提升大模型的空间推理能力,无需标注数据。
The Art of Interrogation: Consistency Amplifies Factuality in Spatial Reasoning

- 用图像和文本变换构建自监督奖励机制,强化几何与语义一致性。
- 在多个空间推理任务上达到接近有监督训练的准确率。
- 适合研究大模型推理对齐与零样本迁移的学者使用。
当前大型推理模型(LRMs)虽具备强大泛化能力,但在空间推理任务中表现不佳。现有方法将此差距归因于知识缺失,依赖监督微调(SFT)从外部视觉源或合成引擎获取标注数据。本文认为,许多任务的空间推理能力已存在于预训练的LRMs中,但需通过几何2D/3D约束下的逻辑连贯性进行对齐。为此,我们提出一种无需真实标注的自监督强化学习框架,通过形式化一致性验证器——即在变换下检查几何与语义一致性的奖励函数——使模型自我提升。我们采用图像翻转及问题中物体顺序交换等文本变换,并设计基于最优传输的强化学习策略OT-GRPO(GRPO的最小匹配变体),专为成对验证器优化。实验表明,该无标签一致性训练方法在准确性上接近有监督训练,且在多样化任务与数据域间保持良好泛化能力。
原文摘要 · Abstract (English)
Current Large Reasoning Models (LRMs) exhibit remarkable general capabilities but significantly underperform in spatial reasoning tasks. Existing approaches treat this gap as a knowledge deficit, relying on supervised fine-tuning (SFT) to ingest labeled spatial data from external vision sources or synthetic engines. In contrast, we argue that for many tasks, spatial reasoning capabilities are already present in pre-trained LRMs but require alignment through logical coherence under geometric 2D and 3D constraints. In this work, we propose a self-supervised reinforcement learning (RL) framework that targets the internal reasoning process without requiring ground-truth annotations. By formalizing the notion of consistency verifiers -- reward functions that check for geometric and semantic consistency under transformations -- we demonstrate that models can improve their spatial reasoning abilities. We use both image transformations, like flipping, and textual transformations, like swapping the order of objects in the question, and propose a new optimal transport-based RL strategy, OT-GRPO, which is a minimal-matching variant of group relative policy optimization tailored to pairwise verifiers. We show that this label-free consistency training approaches the accuracy of models trained with ground-truth supervision and achieves similar generalization across diverse tasks and data domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。