通过视图一致奖励优化,提升多模态模型的空间推理能力。
SVQA-R1: Reinforcing Spatial Reasoning in MLLMs via View-Consistent Reward Optimization
- 设计视图一致奖励机制,通过物体空间关系扰动增强推理
- 在多个空间视觉问答数据集上显著提升准确率
- 无需监督微调即可生成可解释的推理路径,适合空间理解研究
空间推理仍是现有视觉-语言模型中未充分发展的能力,尤其在需要理解相对位置、距离和物体配置的空间视觉问答(Spatial VQA)任务中。受 DeepSeek-R1 中基于规则的强化学习(RL)提升语言模型推理能力的启发,我们提出 SVQA-R1,首个将 R1 风格训练扩展至空间 VQA 的框架。具体地,我们引入 Spatial-GRPO——一种新的分组式强化学习策略,通过扰动物体间空间关系(如镜像翻转)构建视图一致奖励,促使模型建立一致且具象的空间理解。我们的模型在空间 VQA 基准测试中实现显著更高的准确率,并在未使用监督微调数据的情况下展现出可解释的推理路径。大量实验与可视化验证了 SVQA-R1 在多个空间推理基准上的有效性。
原文摘要 · Abstract (English)
Spatial reasoning remains a critical yet underdeveloped capability in existing vision-language models (VLMs), especially for Spatial Visual Question Answering (Spatial VQA) tasks that require understanding relative positions, distances, and object configurations. Inspired by the R1 paradigm introduced in DeepSeek-R1, which enhances reasoning in language models through rule-based reinforcement learning (RL), we propose SVQA-R1, the first framework to extend R1-style training to spatial VQA. In particular, we introduce Spatial-GRPO, a novel group-wise RL strategy that constructs view-consistent rewards by perturbing spatial relations between objects, e.g., mirror flipping, thereby encouraging the model to develop a consistent and grounded understanding of space. Our model, SVQA-R1, not only achieves dramatically improved accuracy on spatial VQA benchmarks but also exhibits interpretable reasoning paths even without using supervised fine-tuning (SFT) data. Extensive experiments and visualization demonstrate the effectiveness of SVQA-R1 across multiple spatial reasoning benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。