arXiv:2607.28285cs.CV2026-07中稿 · ACM MM 2026

用详细长描述提升单目深度估计在复杂场景下的鲁棒性

Beyond Visual Ambiguity: Guiding Robust Monocular Depth Estimation in Challenging Scenarios via Detailed Long Captions

论文配图:Beyond Visual Ambiguity: Guiding Robust Monocular Depth Estimation in Challenging Scenarios via Detailed Long Captions
图 1 · 摘自论文原文
  • 通过长文本描述捕捉物体间空间关系,增强语义引导
  • 动态编码器提取细粒度文本特征,自适应解码器提升深度精度
  • 特别适合处理非朗伯表面和恶劣天气等挑战场景

单目深度估计因单图像信息有限,易受非朗伯表面和恶劣天气带来的视觉模糊影响。现有方法多通过图像修复或增强单独应对,效果有限。语言作为视觉的互补模态,可借助详细长描述增强视觉-语言模型的感知能力。但此前语言融合的单目深度估计方法受限于短文本输入、粗粒度全局特征学习及解码阶段语言引导不足。为此,本文提出CapDepth框架,利用详细长描述缓解复杂场景下的视觉模糊问题。首先设计包含多原子句空间关系的长文本输入模板;其次引入动态文本编码器,通过渐进式掩码注意力提取细粒度深度相关特征;最后提出文本自适应解码器,基于稳定自适应层归一化实现文本引导的深度解码。大量实验验证其有效性:在非朗伯表面场景下深度误差降低25.0%,恶劣天气下降低22.0%,显著优于现有方法。

原文摘要 · Abstract (English)

Monocular depth estimation (MDE) faces challenges with non-Lambertian surfaces and adverse weather conditions due to the visual ambiguities inherent in single-image limited information. Existing works address them in isolation via image inpainting or augmentation, yielding limited robustness gains. Language, as a powerful complementary modality to vision, is demonstrated to enhance the visual perception capabilities of vision-language models (VLMs) via detailed long captions. However, prior language-integrated MDE methods fail to fully harness this potential due to short text input with limited information, coarse global text feature learning, and limited language guidance during depth decoding. To address these limitations, we propose CapDepth, a novel framework for robust MDE that leverages guidance from detailed long captions to alleviate visual ambiguities in both challenging scenarios. First, we design a detailed long caption input template that explicitly conveys rich spatial relationships among multiple atom sentences. Second, a dynamic caption encoder is introduced to extract fine-grained depth-relevant text features via progressive masked attention. Finally, we propose a text-adaptive decoder that guides enhanced depth decoding with text features via stable adaptive layer normalization. Extensive experiments validate the efficacy of CapDepth, which outperforms state-of-the-art methods, achieving depth error reductions of 25.0% on non-Lambertian surfaces and 22.0% under adverse weather conditions.

单目深度视觉语言长文本鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。