arXiv:2609.03729cs.CV2026-09

将视觉语言模型的空间推理分解为四个维度,提升对三维世界的理解能力。

Unfold The World: Factorize 4D Properties in Reinforcing Spatial Reasoning

论文配图:Unfold The World: Factorize 4D Properties in Reinforcing Spatial Reasoning
图 1 · 摘自论文原文
  • 把空间推理拆解为平面对应、深度一致、时间可逆三个可验证目标。
  • 在多视角和视频基准上分别提升5.9%和4.5%的4D推理性能。
  • 适合关注视觉语言模型物理世界理解的科研人员与工程师。

尽管视觉语言模型(VLMs)在通用多模态任务中表现卓越,但在推理物理世界时仍本质上是“扁平”的。我们指出,这一空间瓶颈源于深刻的维度错配:VLMs 被训练用于解读二维投影,而真正的空间推理需要恢复隐含的三维几何与时间连续性。为应对这一高维复杂性,我们倡导从单一学习转向“分而治之”范式。本文提出 FactoSR,一种因子化强化学习框架,显式解析视觉投影所折叠的维度。其核心是将统一的世界一致性推理问题分解为三个正交的几何子目标:平面对应(XY)、深度一致性(Z)与时间可逆性(T)。通过在一个统一策略学习机制中优化这些可验证约束,我们有效将病态的投影恢复问题转化为一系列可操作的推理步骤。在多视角与视频基准上的大量评估表明,这种优雅的分解显著提升了3D与4D推理能力,在 VSI-Bench 上取得5.9%的提升,在 All-Angles-Bench 上取得4.5%的提升。研究结果表明,强化显式的因子化4D一致性是推动VLMs演变为鲁棒、世界感知推理者的关键一步。

原文摘要 · Abstract (English)

Despite the remarkable prowess of Vision-Language Models (VLMs) in general multimodal tasks, they remain fundamentally ``flat'' when reasoning about the physical world. We argue that this spatial bottleneck stems from a profound dimensional mismatch: while VLMs are trained to interpret 2D projections, true spatial reasoning demands the recovery of latent 3D geometry and temporal continuity. To conquer this high-dimensional complexity, we advocate a shift from monolithic learning to a ``divide and conquer'' paradigm. We present FactoSR, a factorized reinforcement learning framework that explicitly interpret the dimensions collapsed by visual projection. At its core, FactoSR decomposes the monolithic problem of world-consistent reasoning into three orthogonal, geometric sub-objectives: planar correspondence ($XY$), depth consistency ($Z$), and temporal reversibility ($T$). By optimizing these verifiable constraints within a unified policy learning mechanism, we effectively transform an ill-posed projection recovery problem into a series of tangible reasoning steps. Extensive evaluations on multi-view and video benchmarks demonstrate that this elegant decomposition yields substantial gains in 3D and 4D reasoning, achieving a 5.9% boost on VSI-Bench and 4.5% on All-Angles-Bench. Our findings suggest that reinforcing explicit, factorized 4D consistency is a critical step toward evolving VLMs into robust, world-aware reasoners.

空间推理多模态强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。