用视觉信息消除语言歧义,让单目深度估计更准更稳定
VGLD: Visually-Guided Linguistic Disambiguation for Monocular Depth Scale Recovery
- 结合图像与文本,通过视觉引导消除语言描述歧义
- 在NYUv2和KITTI上显著降低深度尺度偏差,提升准确性
- 可作为通用轻量模块,零样本下仍保持良好性能
单目深度估计分为相对深度(无真实尺度)和度量深度(具真实尺度)两类。相对方法虽灵活高效,但缺乏实际尺度限制了下游应用。一种有前景的方案是从文本描述中推断绝对尺度,但语言描述易受视角和风格影响而产生歧义。为此,我们提出VGLD(Visually-Guided Linguistic Disambiguation)框架,利用高层视觉语义解决文本输入的歧义问题。通过联合编码图像与文本,VGLD预测一组全局线性变换参数,将相对深度图对齐至度量尺度。该视觉引导的消歧策略提升了尺度估计的稳定性与精度。我们在代表性模型MiDaS和DepthAnything上,基于标准室内(NYUv2)和室外(KITTI)基准进行评估。结果表明,VGLD能有效缓解由不一致或模糊语言导致的尺度偏差,实现鲁棒且准确的度量预测。此外,多数据集训练后,VGLD可作为通用轻量对齐模块,在零样本设置下仍保持强性能。
原文摘要 · Abstract (English)
Monocular depth estimation can be broadly categorized into two directions: relative depth estimation, which predicts normalized or inverse depth without absolute scale, and metric depth estimation, which aims to recover depth with real-world scale. While relative methods are flexible and data-efficient, their lack of metric scale limits their utility in downstream tasks. A promising solution is to infer absolute scale from textual descriptions. However, such language-based recovery is highly sensitive to natural language ambiguity, as the same image may be described differently across perspectives and styles. To address this, we introduce VGLD (Visually-Guided Linguistic Disambiguation), a framework that incorporates high-level visual semantics to resolve ambiguity in textual inputs. By jointly encoding both image and text, VGLD predicts a set of global linear transformation parameters that align relative depth maps with metric scale. This visually grounded disambiguation improves the stability and accuracy of scale estimation. We evaluate VGLD on representative models, including MiDaS and DepthAnything, using standard indoor (NYUv2) and outdoor (KITTI) benchmarks. Results show that VGLD significantly mitigates scale estimation bias caused by inconsistent or ambiguous language, achieving robust and accurate metric predictions. Moreover, when trained on multiple datasets, VGLD functions as a universal and lightweight alignment module, maintaining strong performance even in zero-shot settings. Code will be released upon acceptance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。