用语言估计深度范围,视觉特征精调,实现单目深度度量恢复
Language as Prior, Vision as Calibration: Metric Scale Recovery for Monocular Depth Estimation
- 用文本生成不确定度约束的深度校准范围,避免直接依赖文本点估计
- 在NYUv2和KITTI上提升域内精度,零样本迁移至SUN-RGBD和DDAD表现更鲁棒
- 仅训练轻量校准头,保持主干模型冻结,适合资源受限场景
相对深度基础模型泛化能力强,但单目度量深度因全局尺度不可识别且对域偏移敏感仍属病态问题。在固定主干的校准设置下,我们通过逆深度空间中的图像特定仿射变换恢复度量深度,并仅训练轻量级校准头,保持相对深度主干和CLIP文本编码器不变。由于描述文本提供粗粒度但噪声大的尺度线索,且受表述方式与缺失物体影响,我们不采用文本单一点估计,而是利用语言预测一个不确定性感知的可行校准参数包络线,覆盖无约束空间;再结合池化后的多尺度冻结视觉特征,从中选取图像特定的校准参数。训练时,逆深度空间的闭式最小二乘法提供每张图像的监督信号,用于学习该包络线及所选校准。在NYUv2和KITTI上的实验显示域内精度提升,零样本迁移至SUN-RGBD和DDAD也优于强基线语言模型。
原文摘要 · Abstract (English)
Relative-depth foundation models transfer well, yet monocular metric depth remains ill-posed due to unidentifiable global scale and heightened domain-shift sensitivity. Under a frozen-backbone calibration setting, we recover metric depth via an image-specific affine transform in inverse depth and train only lightweight calibration heads while keeping the relative-depth backbone and the CLIP text encoder fixed. Since captions provide coarse but noisy scale cues that vary with phrasing and missing objects, we use language to predict an uncertainty-aware envelope that bounds feasible calibration parameters in an unconstrained space, rather than committing to a text-only point estimate. We then use pooled multi-scale frozen visual features to select an image-specific calibration within this envelope. During training, a closed-form least-squares oracle in inverse depth provides per-image supervision for learning the envelope and the selected calibration. Experiments on NYUv2 and KITTI improve in-domain accuracy, while zero-shot transfer to SUN-RGBD and DDAD demonstrates improved robustness over strong language-only baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。