提出首个全景度量3D重建框架,解决无序采集下的全局漂移问题。
Argus: Metric Panoramic 3D Reconstruction for Indoor Scenes

- 用可学习的共视模块选择最优参考视图,避免坐标锚点导致的误差累积。
- 在10000场景数据集上实现领先精度,深度估计误差低于2.5%。
- 适合需要高精度室内三维重建的研究者和工业应用开发者。
由于缺乏大规模全景RGB-D训练数据,全景数据的度量前馈3D重建仍研究不足。我们提出了Realsee3D,一个包含10,000个室内场景(1,000个真实、9,000个合成)的混合数据集,包含299,000个全景视角及精确度量标注,并基于此训练了Argus——一个用于度量全景3D重建的前馈网络。在Realsee3D的稀疏无序采集设定下,不当的坐标锚点会导致全局位姿漂移。Argus通过学习的共视模块,选择几何最优参考视图以锚定度量世界坐标系。为进一步提升多任务学习效果,我们将双向像素到世界映射分解为可解释的子步骤,施加逐步监督与跨坐标联合约束,强化预测分支间的几何一致性。在Realsee3D基准测试中,Argus在相机位姿估计、深度估计和点云重建任务上均达到当前最优度量性能。
原文摘要 · Abstract (English)
Metric feed-forward 3D reconstruction for panoramic data remains under-explored due to the lack of large-scale panoramic RGB-D training data. We present Realsee3D, a hybrid dataset of 10K indoor scenes (1K real, 9K synthetic) with 299K panoramic viewpoints and precise metric annotations, and Argus, a feed-forward network trained on it for metric panoramic 3D reconstruction. In the sparse unordered capture setting of Realsee3D, a poorly chosen coordinate anchor can cause global pose drift. Argus addresses this with a learned covisibility module that selects the geometrically optimal reference view to anchor the metric world frame. To further improve multi-task learning, we decompose the bidirectional pixel-to-world mapping into interpretable sub-steps with per-step supervision and cross-coordinate joint constraints, reinforcing geometric consistency across prediction branches. On the Realsee3D benchmark, Argus achieves state-of-the-art metric performance in camera pose estimation, depth estimation, and point cloud reconstruction. Project page: https://argus-paper.realsee.ai.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。