让AI在空间推理时能可视化中间状态,提升多步决策的可靠性。
SVoT: State-aware Visualization-of-Thought for Spatial Reasoning via Reinforcement Learning

- 通过强化学习生成可验证的视觉化中间状态和推理链。
- 在新构建的5个复杂场景中,准确率提升最高达65%。
- 适合研究多步空间推理与可信AI决策的学者使用。
空间推理仍是多模态大模型的难点,因其需对中间状态及状态转移进行可靠多跳推断。现有方法常忽略中间状态验证,将状态转移视为隐式过程,限制了多跳推理的可靠性。为此,我们提出状态感知的可视化思维框架SVoT,基于强化学习生成交错的可验证中间状态与可视化内容。SVoT将转移推理链融入生成过程,使模型通过交错的文本与视觉推理验证动作前提与效果。我们采用分组相对策略优化(GRPO)训练,并通过奖励设计实现验证机制,评估不同细粒度奖励的有效性。由于现有基准将状态转移简化为单变量更新,大幅降低问题难度,我们扩展经典环境并引入两个新领域Pacman与Gather,要求多对象交互与数值推理。这些领域支持对多跳空间推理的系统性评估,包含生成中间状态的定量验证与转移推理检验。搭载转移感知监督的SVoT在所提领域中达到当前最优性能,于分布外测试集上准确率提升最高达65%。
原文摘要 · Abstract (English)
Spatial reasoning remains a challenge for Multimodal Large Language Models (MLLMs), as it requires reliable multi-hop inference over both intermediate states and state transitions. Current studies often leave intermediate states unverified and treat state transitions as implicit processes, which limits reliability in multi-hop spatial reasoning. To address this, we propose State-aware Visualization-of-Thought (SVoT), a reinforcement learning framework that generates interleaved, verifiable intermediate states and visualizations. SVoT integrates transition reasoning chains into the generation processes, enabling the model to verify action preconditions and effects through interleaved textual and visual reasoning. We train SVoT via Group Relative Policy Optimization (GRPO), instantiating verification through reward design and evaluating the efficacy of different fine-grained rewards. As existing benchmarks reduce state transitions to single-variable updates, substantially simplifying the problems, we establish five domains by extending classical environments and introducing two novel domains, Pacman and Gather, that require multi-object interactions and numerical reasoning. These domains support systematic evaluation of multi-hop spatial reasoning with quantitative verification of generated intermediate states and transition reasoning. SVoT with transition-aware supervision achieves state-of-the-art performance across the introduced domains, yielding up to a 65% absolute accuracy gain on out-of-distribution test sets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。