arXiv:2608.12220cs.CVcs.AI2026-08

通过结构化思维链与多目标奖励,提升视觉语言模型的空间推理能力。

SCOUT: Unlocking Enhanced Spatial Reasoning via Structured Chain-of-Thought and Multi-Objective Process Reward

论文配图:SCOUT: Unlocking Enhanced Spatial Reasoning via Structured Chain-of-Thought and Multi-Objective Process Reward
图 1 · 摘自论文原文
  • 设计结构化思维链框架,显式建模三维环境感知。
  • 引入多目标过程奖励机制,实现推理步骤的精准信用分配。
  • 在单图训练下仍可泛化到多图与视频场景,适合空间理解任务。

现有视觉语言模型在鲁棒空间推理方面存在明显瓶颈。尽管近期强化学习方法已取得可验证成果,但其在中间推理步骤间的信用分配表现不佳。同时,结构化推理方法忽视了全面3D理解所需的深度感知。为此,我们提出SCOUT(Structured Chain-Of-Thought Utilizing Process-Supervised RL Training)。具体地,设计了一种结构化思维链(CoT)框架,显式建模3D环境感知以保障稳健的空间理解与推理。此外,提出一种新型强化学习算法,包含多目标过程奖励与定制化的优势估计方法,实现推理轨迹各阶段的细粒度信用分配。为支持该框架,我们构建了SCOUT-24k——一个通过定制化流水线生成的结构化空间推理CoT数据集。大量实验表明,SCOUT-3B在通用空间基准和复杂空间推理任务上分别优于基线模型16.85%和6.3%。值得注意的是,更大规模的SCOUT-7B甚至在性能上超过GPT-4o达4.28%。此外,尽管仅在单图数据上训练,SCOUT-7B仍展现出对多图与视频场景的强大跨域泛化能力。这些实证结果使SCOUT成为迈向下一代空间感知视觉语言模型的关键一步。

原文摘要 · Abstract (English)

Existing Vision-Language Models (VLMs) exhibits a critical bottleneck in robust spatial reasoning. Recent reinforcement learning (RL) methods aim to close this gap with verifiable outcomes, yet they suffer from poor credit assignment across intermediate reasoning steps. Concurrently, structured reasoning approaches overlook the critical depth perception necessary for comprehensive 3D understanding. To address these challenges, we propose SCOUT (Structured Chain-Of-Thought Utilizing Process-Supervised RL Training). Specifically, we design a structured Chain-of-Thought (CoT) framework that explicitly models 3D environmental perception to ensure robust spatial understanding and reasoning. Furthermore, we introduce a novel RL algorithm featuring multi-objective process rewards and a tailored advantage estimation method, facilitating fine-grained credit assignment across distinct segments of the reasoning trajectory. To support our framework, we develop SCOUT-24k, a structured spatial reasoning CoT dataset synthesized through a customized pipeline. Extensive evaluations demonstrate that SCOUT-3B improves upon baseline models by 16.85% and 6.3% on general spatial benchmarks and complex spatial reasoning tasks respectively. Notably, our larger SCOUT-7B even outperforms GPT-4o by a margin of 4.28%. Moreover, despite being trained exclusively on single image, SCOUT-7B exhibits robust out-of-domain generalization to multi-image and video scenarios. These empirical results render SCOUT as a critical step towards next generation of spatially-aware VLMs.

空间推理思维链强化学习视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。