arXiv:2511.00710cs.AI2025-11ACL

RLVR让视觉语言模型突破空间推理极限,实现在原模型完全失败的任务上成功导航。

Does RLVR Extend Reasoning Boundaries? Investigating Capability Expansion in Vision-Language Models

  • 用可调控的合成迷宫框架测试模型,精确控制路径长度和转弯数以评估推理难度。
  • 在基线模型0%准确率的情况下,优化后模型在复杂迷宫中实现成功导航,突破能力边界。
  • 零样本迁移至真实地图任务仍表现提升,证明是真正推理能力增强而非单纯采样优化。

近期研究认为强化学习结合可验证奖励(RLVR)主要放大预训练分布内的行为,而非引入新能力,但这些结论多集中于纯语言领域,对以视觉为中心的空间推理机制仍缺乏探索。为考察RLVR对视觉语言模型(VLMs)能力边界的影响力,我们提出 extbf{Ariadne}——一个基于合成迷宫导航的可控框架,通过路径长度与转弯数精确调控推理难度。实验表明,应用RLVR可显著扩展空间推理边界:在基线模型始终达到0%准确率、且增加pass@k采样预算也无法改善的情况下,优化后的策略仍能成功导航,说明其有效探索了原分布不可达的搜索空间。此外,尽管训练仅限于合成迷宫,模型在两个真实世界导航基准(MapBench与ReasonMap)上实现零样本性能提升,表明该能力扩张具有实际泛化性,非仅限于采样效率改进。

原文摘要 · Abstract (English)

Recent studies posit that Reinforcement Learning with Verifiable Rewards (RLVR) primarily amplifies behaviors inherent to the pre-training distribution rather than inducing new capabilities, but these insights are predominantly limited to language-only domains, leaving the dynamics of visual-centric spatial reasoning under-explored. To examine the impact of RLVR on the capability boundaries of Vision-Language Models (VLMs), we introduce \textbf{Ariadne}, a controlled framework based on synthetic maze navigation where the reasoning difficulty is precisely regulated by path length and the number of turns. We demonstrate that applying RLVR extends the spatial reasoning boundary, achieving success on problems where the base policy VLM consistently attains $0\%$ accuracy despite increasing pass@k sampling budgets, indicating that the optimized policy successfully navigates search spaces that were effectively unreachable by the base distribution. Furthermore, despite being trained exclusively on synthetic mazes, we evaluate the model on two real-world navigation benchmarks (MapBench and ReasonMap) in a zero-shot setting. The observed improvements in these out-of-domain tasks suggest genuine spatial reasoning capability expansion rather than mere sampling efficiency.

视觉语言模型空间推理强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。