将不确定性量化与大模型结合,提升单目深度估计的可靠性。
A Critical Synthesis of Uncertainty Quantification and Foundation Models in Monocular Depth Estimation
- 融合五种不确定性方法到DepthAnythingV2模型中
- 高斯负对数似然损失微调效果最佳,性能不降且效率高
- 为深度估计可解释性提供新思路,适合安全关键场景
尽管近期基础模型在单目深度估计中取得显著进展,但其在真实世界中安全可靠部署的路径仍不明确。绝对距离预测任务尤其困难,即使最先进的基础模型也易出现严重误差。由于不确定性量化被视为解决此类问题的有前景方向,本文将五种不同的不确定性量化方法与当前最优的DepthAnythingV2基础模型相结合。为覆盖广泛的距离范围,我们在四个多样化数据集上评估其表现。结果表明,使用高斯负对数似然损失(GNLL)进行微调是一种特别有前景的方法,在保持预测性能和计算效率(训练与推理时间均与基线相当)的同时,提供可靠的不确定性估计。通过在单目深度估计中融合不确定性量化与基础模型,本文为未来提升模型性能与可解释性的研究奠定关键基础。将这一综合分析拓展至语义分割、位姿估计等关键任务,有望推动更安全、更可靠的机器视觉系统发展。
原文摘要 · Abstract (English)
While recent foundation models have enabled significant breakthroughs in monocular depth estimation, a clear path towards safe and reliable deployment in the real-world remains elusive. Metric depth estimation, which involves predicting absolute distances, poses particular challenges, as even the most advanced foundation models remain prone to critical errors. Since quantifying the uncertainty has emerged as a promising endeavor to address these limitations and enable trustworthy deployment, we fuse five different uncertainty quantification methods with the current state-of-the-art DepthAnythingV2 foundation model. To cover a wide range of metric depth domains, we evaluate their performance on four diverse datasets. Our findings identify fine-tuning with the Gaussian Negative Log-Likelihood Loss (GNLL) as a particularly promising approach, offering reliable uncertainty estimates while maintaining predictive performance and computational efficiency on par with the baseline, encompassing both training and inference time. By fusing uncertainty quantification and foundation models within the context of monocular depth estimation, this paper lays a critical foundation for future research aimed at improving not only model performance but also its explainability. Extending this critical synthesis of uncertainty quantification and foundation models into other crucial tasks, such as semantic segmentation and pose estimation, presents exciting opportunities for safer and more reliable machine vision systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。