将相机物理特性融入模型,实现单目深度实时估算
Vision-Language Embodiment for Monocular Depth Estimation
- 把相机内在参数嵌入网络,通过与环境交互计算深度
- 结合图像与文本描述,提升对场景几何和语义的理解
- 适合需要动态感知的机器人视觉系统
单目深度估计是机器人感知与视觉任务的核心问题,但仅凭单张图像进行3D重建存在固有不确定性。现有模型主要依赖图像间关系进行监督训练,常忽略相机自身的内在信息。本文提出一种方法,将相机模型及其物理特性嵌入深度学习模型,通过与道路环境的实时交互,仅利用相机固有属性即可实时计算场景深度,无需额外设备。结合图像特征与文本描述中的环境内容和深度先验,模型可同时获取几何与视觉细节。该多模态融合策略整合了图像与语言这两种固有模糊模态的互补优势,实现更鲁棒的单目深度估计。实验证明,该方法在不同场景下均提升了模型性能。
原文摘要 · Abstract (English)
Depth estimation is a core problem in robotic perception and vision tasks, but 3D reconstruction from a single image presents inherent uncertainties. Current depth estimation models primarily rely on inter-image relationships for supervised training, often overlooking the intrinsic information provided by the camera itself. We propose a method that embodies the camera model and its physical characteristics into a deep learning model, computing embodied scene depth through real-time interactions with road environments. The model can calculate embodied scene depth in real-time based on immediate environmental changes using only the intrinsic properties of the camera, without any additional equipment. By combining embodied scene depth with RGB image features, the model gains a comprehensive perspective on both geometric and visual details. Additionally, we incorporate text descriptions containing environmental content and depth information as priors for scene understanding, enriching the model's perception of objects. This integration of image and language - two inherently ambiguous modalities - leverages their complementary strengths for monocular depth estimation. The real-time nature of the embodied language and depth prior model ensures that the model can continuously adjust its perception and behavior in dynamic environments. Experimental results show that the embodied depth estimation method enhances model performance across different scenes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。