解决单目几何估计中远距离物体尺度低估问题,构建真实世界大尺度数据集。
Honey, I Shrunk the Arc de Triomphe!

- 构建跨场景真实世界数据集MetricScenes,融合网络图片与立体图像。
- 利用地理标签和相机基线恢复绝对尺度,深度图精度显著提升。
- 在开放场景中显著缓解尺度坍缩,适合高精度3D重建研究者使用。
单目几何估计在大规模数据聚合下取得进展,但现有基础模型普遍存在尺度坍缩问题:远距离地标和广阔景观的度量估计严重偏低。我们推测该性能差距源于训练数据瓶颈——现有度量数据集受限于硬件,多为车载激光雷达或短距离室内扫描,或为缺乏物理世界语义复杂性的合成数据。为此,我们构建了一个新的、真实世界的度量基准数据集MetricScenes,数据来源包括互联网图片集合与立体图像。通过现成方法估计相机位姿与初始深度图,并利用地理标记元数据及已知立体相机基线恢复绝对尺度。同时,提出一种两阶段泊松补全方法,提升MetricScenes生成的深度图质量。在该数据集上微调MoGe-2,显著缓解了尺度坍缩,在非约束、开放域场景中实现更优的度量精度,同时保持标准基准上的顶尖性能。
原文摘要 · Abstract (English)
Metric scale monocular geometry estimation has seen significant progress through large-scale data aggregation, yet current foundation models suffer from a persistent ''scale-collapse'' phenomenon: distant landmarks and vast landscapes are metrically underestimated. We hypothesize that this performance gap stems from a training data bottleneck, where existing metric-scale datasets are hardware-constrained to homogenous vehicle-captured LiDAR or short-range indoor scans, or consist of synthetic data that lacks the semantic complexity of the physical world. To bridge this gap, we curate a new metrically-grounded, in-the-wild dataset that we call MetricScenes, gathered from a variety of sources including Internet photo collections and stereo imagery. We estimate camera poses and initial depth maps for each scene using off-the-shelf methods, and recover absolute scale from geo-tagged metadata as well as known stereo camera baselines. We also improve the quality of depth maps derived from MetricScenes via a new two-stage Poisson completion method. Fine-tuning MoGe-2 on our dataset significantly mitigates scale-collapse and achieves superior metric accuracy in unconstrained, open-domain scenes while maintaining state-of-the-art performance on standard benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。