arXiv:2512.07733cs.CV2025-12被引 11

让AI像人一样主动想象空间关系,提升复杂场景推理能力。

SpatialDreamer: Incentivizing Spatial Reasoning via Active Mental Imagery

  • 通过强化学习构建主动探索与视觉想象的闭环推理流程。
  • 在多个基准测试中表现优异,尤其在长链条空间推理任务上超越现有方法。
  • 适合研究具身智能、空间认知与多模态模型的学者参考。

尽管多模态大语言模型(MLLMs)在场景理解方面取得进展,其在需要心理模拟的复杂空间推理任务上的表现仍显著受限。现有方法多依赖对空间数据的被动观察,未能内化主动的心理意象过程。为此,我们提出SpatialDreamer,一种基于强化学习的框架,通过主动探索、世界模型驱动的视觉想象以及基于证据的推理构成闭环过程。为解决长链条推理任务中缺乏细粒度奖励监督的问题,我们提出几何策略优化(GeoPO),引入树状采样和带几何一致性约束的逐步奖励估计。大量实验表明,SpatialDreamer在多个挑战性基准上达到具有竞争力的性能,标志着MLLMs在类人主动空间心理模拟方面的关键进展。

原文摘要 · Abstract (English)

Despite advancements in Multi-modal Large Language Models (MLLMs) for scene understanding, their performance on complex spatial reasoning tasks requiring mental simulation remains significantly limited. Current methods often rely on passive observation of spatial data, failing to internalize an active mental imagery process. To bridge this gap, we propose SpatialDreamer, a reinforcement learning framework that enables spatial reasoning through a closedloop process of active exploration, visual imagination via a world model, and evidence-grounded reasoning. To address the lack of fine-grained reward supervision in longhorizontal reasoning tasks, we propose Geometric Policy Optimization (GeoPO), which introduces tree-structured sampling and step-level reward estimation with geometric consistency constraints. Extensive experiments demonstrate that SpatialDreamer delivers highly competitive results across multiple challenging benchmarks, signifying a critical advancement in human-like active spatial mental simulation for MLLMs.

空间推理强化学习心理意象多模态模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。