让3D重建模型学会利用多相机硬件结构,提升精度与鲁棒性。
Rig3R: Rig-Aware Conditioning for Learned 3D Reconstruction
- 基于相机阵列结构设计条件感知的隐空间,可处理有无元数据的情况
- 在多个真实场景数据集上,3D重建与姿态估计指标领先17-45% mAA
- 一次前向传播完成重建、姿态估计和阵列结构推断,无需迭代优化
从多摄像头阵列中估计智能体姿态与三维场景结构是具身AI应用(如自动驾驶)的核心任务。现有学习方法如DUSt3R在多视角设置下表现优异,但将图像视为无结构集合,在已知或可推断结构的同步阵列场景中效果受限。为此,我们提出Rig3R,一种可融入阵列结构信息的通用多视角重建模型,支持在有/无元数据时自动学习并利用阵列结构。Rig3R通过相机编号、时间戳和阵列位姿等可选元数据,构建对阵列敏感的潜在空间,联合预测点图和两类射线图:相对于全局坐标系的姿态射线图,以及随时间一致的阵列中心坐标系下的阵列射线图。当缺乏元数据时,阵列射线图可直接从输入图像推断出阵列结构。Rig3R在3D重建、相机姿态估计与阵列发现任务上均达到当前最优,跨多样化真实世界阵列数据集,相比传统与学习方法提升17-45% mAA,且仅需单次前向传播,无需后处理或迭代优化。
原文摘要 · Abstract (English)
Estimating agent pose and 3D scene structure from multi-camera rigs is a central task in embodied AI applications such as autonomous driving. Recent learned approaches such as DUSt3R have shown impressive performance in multiview settings. However, these models treat images as unstructured collections, limiting effectiveness in scenarios where frames are captured from synchronized rigs with known or inferable structure. To this end, we introduce Rig3R, a generalization of prior multiview reconstruction models that incorporates rig structure when available, and learns to infer it when not. Rig3R conditions on optional rig metadata including camera ID, time, and rig poses to develop a rig-aware latent space that remains robust to missing information. It jointly predicts pointmaps and two types of raymaps: a pose raymap relative to a global frame, and a rig raymap relative to a rig-centric frame consistent across time. Rig raymaps allow the model to infer rig structure directly from input images when metadata is missing. Rig3R achieves state-of-the-art performance in 3D reconstruction, camera pose estimation, and rig discovery, outperforming both traditional and learned methods by 17-45% mAA across diverse real-world rig datasets, all in a single forward pass without post-processing or iterative refinement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。