提升3D视觉几何估计性能,关键在数据质量与损失设计。
Unlocking the Power of Critical Factors for 3D Visual Geometry Estimation

- 通过消融实验发现数据多样性是性能提升关键
- 传统置信度损失和梯度损失反而抑制模型表现
- 结合序列与帧级联合监督,适合多帧几何估计任务
前馈式视觉几何估计近期进展迅速,但多帧模型虽具更好跨帧一致性,却常在单帧精度上逊于强帧级方法。本文通过系统消融研究揭示关键因素:1)提升数据多样性和质量可进一步推动现有先进方法性能;2)常用置信度感知损失与基于梯度的损失机制可能无意中阻碍表现;3)通过序列与帧级对齐联合监督可提升结果,而局部区域对齐反而导致性能下降。此外,引入一致性损失函数以对齐深度图、相机参数与点云,并采用高效架构融合高分辨率信息。将这些设计集成至CARVE模型,该模型在点云重建、视频深度估计及相机位姿/内参估计多个基准上均表现出强且稳健的性能。
原文摘要 · Abstract (English)
Feed-forward visual geometry estimation has recently made rapid progress. However, an important gap remains: multi-frame models usually produce better cross-frame consistency, yet they often underperform strong per-frame methods on single-frame accuracy. This observation motivates our systematic investigation into the critical factors driving model performance through rigorous ablation studies, which reveals several key insights: 1) Scaling up data diversity and quality unlocks further performance gains even in state-of-the-art visual geometry estimation methods; 2) Commonly adopted confidence-aware loss and gradient-based loss mechanisms may unintentionally hinder performance; 3) Joint supervision through both per-sequence and per-frame alignment improves results, while local region alignment surprisingly degrades performance. Furthermore, we introduce two enhancements to integrate the advantages of optimization-based methods and high-resolution inputs: a consistency loss function that enforces alignment between depth maps, camera parameters, and point maps, and an efficient architectural design that leverages high-resolution information. We integrate these designs into CARVE, a resolution-enhanced model for feed-forward visual geometry estimation. Experiments on point cloud reconstruction, video depth estimation, and camera pose/intrinsic estimation show that CARVE achieves strong and robust performance across diverse benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。