提出新评测框架BenchDepth,用五项真实任务评估深度模型实用性。
BenchDepth: Are We on the Right Way to Evaluate Depth Foundation Models?
- 设计五项下游代理任务替代传统对齐评估
- 在8个先进深度模型上验证,揭示现有评测偏差
- 适合关注深度模型真实应用效果的研究者
深度估计是计算机视觉中的基础任务,具有广泛的应用。近年来深度学习的发展催生了强大的深度基础模型(DFMs),但其评估仍面临挑战,源于现有评测协议的不一致。传统基准依赖基于对齐的度量,引入偏差,偏爱特定深度表示,阻碍公平比较。本文提出BenchDepth,通过五个精心选择的下游代理任务评估DFMs:深度补全、立体匹配、单目前馈3D场景重建、SLAM和视觉-语言空间理解。与传统方法不同,本方法基于模型在实际应用中的实用价值进行评估,绕过有争议的对齐过程。我们对八个最先进DFMs进行了基准测试,并深入分析关键发现与观察。希望本工作能引发社区对深度模型评估最佳实践的讨论,为未来深度估计研究与进步铺平道路。
原文摘要 · Abstract (English)
Depth estimation is a fundamental task in computer vision with diverse applications. Recent advancements in deep learning have led to powerful depth foundation models (DFMs), yet their evaluation remains challenging due to inconsistencies in existing protocols. Traditional benchmarks rely on alignment-based metrics that introduce biases, favor certain depth representations, and complicate fair comparisons. In this work, we propose BenchDepth, a new benchmark that evaluates DFMs through five carefully selected downstream proxy tasks: depth completion, stereo matching, monocular feed-forward 3D scene reconstruction, SLAM, and vision-language spatial understanding. Unlike conventional evaluation protocols, our approach assesses DFMs based on their practical utility in real-world applications, bypassing problematic alignment procedures. We benchmark eight state-of-the-art DFMs and provide an in-depth analysis of key findings and observations. We hope our work sparks further discussion in the community on best practices for depth model evaluation and paves the way for future research and advancements in depth estimation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。