arXiv:2505.23400cs.CV2025-05被引 1

融合几何与语义信息,提升单目深度估计在复杂场景下的泛化能力。

Bridging Geometric and Semantic Foundation Models for Generalized Monocular Depth Estimation

  • 通过桥接门融合深度与分割模型的互补优势。
  • 在多个数据集上超越现有方法,尤其擅长处理重叠物体和复杂结构。
  • 仅训练桥接门,大幅降低资源消耗,适合快速部署。

我们提出BriGeS,一种有效融合几何与语义信息的单目深度估计方法。核心是桥接门(Bridging Gate),整合深度与分割基础模型的互补优势。通过注意力温度缩放技术,精细调节注意力机制,避免对特定特征过度聚焦,确保在多样化输入下表现均衡。BriGeS利用预训练基础模型,仅训练桥接门,显著降低资源需求与训练时间,同时保持强泛化能力。在多个挑战性数据集上的实验表明,BriGeS在复杂场景下优于现有最先进方法,能有效处理复杂结构与重叠物体。

原文摘要 · Abstract (English)

We present Bridging Geometric and Semantic (BriGeS), an effective method that fuses geometric and semantic information within foundation models to enhance Monocular Depth Estimation (MDE). Central to BriGeS is the Bridging Gate, which integrates the complementary strengths of depth and segmentation foundation models. This integration is further refined by our Attention Temperature Scaling technique. It finely adjusts the focus of the attention mechanisms to prevent over-concentration on specific features, thus ensuring balanced performance across diverse inputs. BriGeS capitalizes on pre-trained foundation models and adopts a strategy that focuses on training only the Bridging Gate. This method significantly reduces resource demands and training time while maintaining the model's ability to generalize effectively. Extensive experiments across multiple challenging datasets demonstrate that BriGeS outperforms state-of-the-art methods in MDE for complex scenes, effectively handling intricate structures and overlapping objects.

深度估计基础模型多模态融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。