用惯性传感器让单目3D模型恢复真实尺度,无需标注数据。
VI3: Grounding Pretrained 3D Foundation Models with Inertial Cues

- 用IMU预积分生成度量运动参考,锚定3D模型输出尺度。
- 在真实与合成航拍数据上实现无监督度量尺度恢复。
- 适配多种3D模型架构,适合移动设备端三维重建场景。
3D基础模型(3DFMs)能从多视角图像中预测相机位姿和稠密深度,展现出强大的零样本泛化能力。然而,由于单目图像无法观测绝对尺度,其预测的尺度通常不准确。大多数设备内置惯性测量单元(IMUs),可自然补充单目相机的尺度信息。本文提出VI3,一种与模型无关的框架,仅利用IMU读数对预训练的3DFM进行度量锚定。VI3通过初始化并预积分IMU获得度量运动参考,再用于恢复3DFM输出的尺度。方法包含适配多种3DFM架构的可调锚定策略。在合成与真实航拍数据集上的实验表明,VI3可在无真值监督下恢复度量尺度,同时保持几何一致性,在运动条件良好时作为精修手段,运动信息不足时则提供强先验。
原文摘要 · Abstract (English)
3D foundation models (3DFMs) excel at predicting camera poses and dense depth from multiple views of a scene, showcasing strong zero-shot generalization. However, as metric scale is not observable from monocular images, their absolute scale predictions are typically inaccurate. Inertial measurement units (IMUs), present in most devices, naturally complement monocular cameras by observing scaled motion. We introduce VI3, a model-agnostic framework that metrically anchors a pretrained 3DFM using only IMU readings. VI3 initializes and preintegrates the IMU to obtain a metric motion reference, which is then used to recover the scale of the 3DFM outputs. Our method includes adaptable anchoring strategies tailored to diverse 3DFM architectures. Experiments on synthetic and real aerial datasets demonstrate that VI3 recovers metric scale without ground-truth supervision while preserving geometric consistency, acting as a fine refinement under well-conditioned motion and as a strong prior when motion is less informative.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。