arXiv:2410.18962cs.CV2024-10ICLR被引 6

用统一模型同时解决定位与视野预测,提升空间感知能力。

Where Am I and What Will I See: An Auto-Regressive Model for Spatial Localization and View Prediction

  • 提出自回归框架GST,联合建模相机位姿与视角生成
  • 在单图下同时估计位姿并预测新视角,性能优于独立任务方法
  • 适合做三维视觉理解、机器人导航的开发者参考

空间智能是机器在时空维度中感知、推理和行动的能力。近年来大规模自回归模型在多种推理任务中展现出强大能力,但在回答‘我在哪?’和‘我将看到什么?’这类基础空间推理问题上仍表现不佳。现有方法通常将两者分开处理,未能捕捉其内在关联。本文提出生成式空间变换器(Generative Spatial Transformer, GST),一种新型自回归框架,可联合完成空间定位与视图预测。该模型从单张图像中同时估计相机位姿,并预测新位姿下的视图,有效连接空间感知与视觉预测。创新的相机标记化方法使模型能以自回归方式学习2D投影与其对应空间视角的联合分布。统一训练范式首次表明,联合优化位姿估计与新视图合成可显著提升两项任务性能,揭示了空间意识与视觉预测之间的内在联系。

原文摘要 · Abstract (English)

Spatial intelligence is the ability of a machine to perceive, reason, and act in three dimensions within space and time. Recent advancements in large-scale auto-regressive models have demonstrated remarkable capabilities across various reasoning tasks. However, these models often struggle with fundamental aspects of spatial reasoning, particularly in answering questions like "Where am I?" and "What will I see?". While some attempts have been done, existing approaches typically treat them as separate tasks, failing to capture their interconnected nature. In this paper, we present Generative Spatial Transformer (GST), a novel auto-regressive framework that jointly addresses spatial localization and view prediction. Our model simultaneously estimates the camera pose from a single image and predicts the view from a new camera pose, effectively bridging the gap between spatial awareness and visual prediction. The proposed innovative camera tokenization method enables the model to learn the joint distribution of 2D projections and their corresponding spatial perspectives in an auto-regressive manner. This unified training paradigm demonstrates that joint optimization of pose estimation and novel view synthesis leads to improved performance in both tasks, for the first time, highlighting the inherent relationship between spatial awareness and visual prediction.

空间感知自回归位姿估计视图预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。