提升视频中3D空间推理能力,解决视觉模型对空间关系理解不足的问题。
SpaceEra++: A Unified Framework Towards 3D Spatial Reasoning in Video

- 提出新采样策略ScenePick,兼顾空间覆盖与语义信息,增强输入质量。
- 设计SpaceAlign机制,结合绝对坐标与相对关系,精准约束物体空间位置。
- 在多个数据集上超越现有方法,适合机器人导航等需要空间理解的任务。
视觉空间理解指从视觉输入中推断物体关系与场景布局的能力,是机器人导航和具身交互等下游任务的基础。然而,预训练视觉-语言模型受限于固有的2D观测带来的空间不确定性,以及缺乏3D空间理解的数据。为此,我们此前在NeurIPS 2025 Spotlight论文中提出了SpaceEra框架,虽取得显著性能提升,但发现其效果受扫描视频输入不足和推理约束较弱的限制。为此,本文将原框架扩展为综合性系统SpaceEra++,涵盖数据构建、模型设计、训练优化与提示推理。具体而言,为缓解输入不足,提出ScenePick帧采样策略,在保证空间覆盖的同时保留物体语义,生成紧凑且全面的场景表征;为进一步增强空间推理,设计SpaceAlign机制,通过联合利用绝对坐标与相对空间关系,施加成对物体约束,使优化过程更贴近空间准确性。大量实验表明,该框架在多个基准上持续优于强基线,消融实验证明各组件独立及协同贡献有效,进一步分析也为未来研究提供指导。
原文摘要 · Abstract (English)
Visual-spatial understanding, defined as the ability to infer object relationships and scene layouts from visual inputs, is fundamental to downstream tasks such as robotic navigation and embodied interaction. However, pre-trained vision-language models (VLMs) remain constrained by spatial uncertainty stemming from inherently 2D observations and by the scarcity of data for 3D spatial understanding. To address these limitations, we proposed a novel framework, SpaceEra, in the NeurIPS 2025 Spotlight paper. Although it achieved significant performance gains, we further observed that its effectiveness is hindered by insufficient input from scanning videos and weak reasoning constraints. To tackle these newly emerged challenges, we extend the original framework into a comprehensive system, termed SpaceEra++, which spans data construction, model design, training optimization, and prompting inference. Specifically, to alleviate input insufficiency, we introduce ScenePick, a frame sampling strategy that balances spatial coverage with object semantics to produce compact yet comprehensive scene representations. In addition, to enhance spatial reasoning, we develop SpaceAlign, which enforces pairwise object constraints by jointly exploiting absolute coordinates and relative spatial relations, thereby aligning optimization with spatial accuracy. Extensive experiments across multiple benchmarks demonstrate consistent improvements over strong baselines, while ablation studies validate both the individual and joint contributions of each component, and further analyses provide guidance for future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。