提升视觉语言模型的空间推理可靠性,让思考过程更稳定、更依赖真实视觉信息。
ProSR: Process-Shaped Spatial Reasoning for Reliable Chain-of-Thought in VLMs

- 通过反事实不变性与尾部漂移惩罚,约束推理过程的视觉依赖性和稳定性。
- 在多个复杂空间推理任务上,准确率提升且推理路径更稳定。
- 适合关注模型可解释性与鲁棒性的研究人员。
可靠的视觉空间推理仍是视觉语言模型的核心瓶颈。现有主流训练范式主要依赖结果对齐或过程模仿,缺乏对推理过程的显式约束,难以保证真实的视觉依赖性和稳定的推理轨迹。本文构建了一个高质量的链式思维(CoT)数据集,覆盖多样化的空间现象,并诊断模型推理过程,发现强化学习优化中存在两类典型过程退化:虚假锚定(Spurious Grounding),绕过视觉证据;尾部不稳定性(Tail Instability),推理后期不确定性异常升高。为此,提出ProSR框架,通过反事实不变性惩罚和尾部漂移惩罚,将优化目标从单一答案正确性扩展至两个过程层面维度:视觉依赖性和轨迹稳定性。在多个复杂及分布外空间推理基准测试中,ProSR不仅提升答案准确率,还生成更稳定、更依赖视觉证据的推理轨迹。
原文摘要 · Abstract (English)
Reliable spatial reasoning remains a core bottleneck for vision-language models (VLMs). Existing mainstream training paradigms for spatial reasoning largely rely on outcome alignment or process imitation, lacking explicit constraints on the reasoning process, and therefore struggle to ensure genuine visual dependence and stable reasoning trajectories. In this paper, we construct a high-quality CoT dataset covering diverse spatial phenomena and diagnose the model's reasoning process, revealing two typical types of process degradation during reinforcement learning optimization: Spurious Grounding, which bypasses visual evidence, and Tail Instability, where uncertainty abnormally rises in the later stage of reasoning. To address these issues, we propose ProSR, a process-shaping optimization framework for spatial reasoning. Through a Counterfactual Invariance Penalty and a Tail Drift Penalty, ProSR extends the optimization objective from single answer correctness to two process-level dimensions: visual dependence and trajectory stability. Experiments on multiple complex and out-of-distribution spatial reasoning benchmarks show that ProSR improves answer accuracy while generating reasoning trajectories that are more stable and more dependent on visual evidence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。