探究大模型如何无监督理解单目深度,发现越大的模型越像人。
DepthCues: Evaluating Monocular Depth Perception in Large Vision Models
- 构建新基准DepthCues,测试模型对人类视觉深度线索的理解能力。
- 20个模型对比显示,更大更新的模型具备类人深度感知能力。
- 无需密集标注数据,用该基准微调即可提升模型深度估计性能。
大规模预训练视觉模型日益普及,其表达能力强且具备泛化性,可助力多种下游任务。近期研究揭示这些模型具备高水平几何理解能力,尤其在深度感知方面。然而,这些模型在未接受显式深度监督的情况下,如何形成深度感知仍不明确。为此,我们考察了类似人类视觉系统的单目深度线索是否在这些模型中涌现。本文提出新基准DepthCues,用于评估深度线索理解能力,并对20个多样且具有代表性的预训练视觉模型进行了分析。结果表明,更近期、更大的模型展现出类人深度线索。此外,通过在DepthCues上微调模型,即使无密集深度标注,也能有效提升深度估计性能。为支持后续研究,本工作将公开基准与评估代码。
原文摘要 · Abstract (English)
Large-scale pre-trained vision models are becoming increasingly prevalent, offering expressive and generalizable visual representations that benefit various downstream tasks. Recent studies on the emergent properties of these models have revealed their high-level geometric understanding, in particular in the context of depth perception. However, it remains unclear how depth perception arises in these models without explicit depth supervision provided during pre-training. To investigate this, we examine whether the monocular depth cues, similar to those used by the human visual system, emerge in these models. We introduce a new benchmark, DepthCues, designed to evaluate depth cue understanding, and present findings across 20 diverse and representative pre-trained vision models. Our analysis shows that human-like depth cues emerge in more recent larger models. We also explore enhancing depth perception in large vision models by fine-tuning on DepthCues, and find that even without dense depth supervision, this improves depth estimation. To support further research, our benchmark and evaluation code will be made publicly available for studying depth perception in vision models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。