arXiv:2512.08700cs.CV2025-12中稿 · ICRA

轻量级模型实现全景单目深度实时估计,解决尺度不准与计算开销大问题。

Scale-invariant and View-relational Representation Learning for Full Surround Monocular Depth

  • 用知识蒸馏将大模型的深度判别能力迁移到小模型,提升精度与效率
  • 引入深度分箱与跨视图关系蒸馏,使预测结果具备真实尺度与跨视角一致性
  • 在DDAD和nuScenes数据集上优于传统方法,满足实时性要求

近期基础模型在单目深度估计中展现出强大泛化能力,但直接用于全景单目深度估计(FSMDE)面临两大挑战:高计算开销限制实时性能,以及难以估计度量尺度深度(因模型通常仅预测相对深度)。为此,我们提出一种新型知识蒸馏策略,将基础模型的鲁棒深度知识迁移至轻量级FSMDE网络。该方法结合传统分类蒸馏与深度分箱模块,构建混合回归框架,通过交叉交互蒸馏,将基础模型的尺度不变深度分箱概率传递给学生网络,并引导其从真实深度中学习度量尺度分箱中心。同时,提出视图关联蒸馏,编码相邻相机视角间的结构关系,增强跨视角深度一致性。在DDAD和nuScenes数据集上的实验表明,本方法显著优于传统监督方法及现有知识蒸馏方案,且在性能与效率间取得良好平衡,满足实时性需求。

原文摘要 · Abstract (English)

Recent foundation models demonstrate strong generalization capabilities in monocular depth estimation. However, directly applying these models to Full Surround Monocular Depth Estimation (FSMDE) presents two major challenges: (1) high computational cost, which limits real-time performance, and (2) difficulty in estimating metric-scale depth, as these models are typically trained to predict only relative depth. To address these limitations, we propose a novel knowledge distillation strategy that transfers robust depth knowledge from a foundation model to a lightweight FSMDE network. Our approach leverages a hybrid regression framework combining the knowledge distillation scheme--traditionally used in classification--with a depth binning module to enhance scale consistency. Specifically, we introduce a cross-interaction knowledge distillation scheme that distills the scale-invariant depth bin probabilities of a foundation model into the student network while guiding it to infer metric-scale depth bin centers from ground-truth depth. Furthermore, we propose view-relational knowledge distillation, which encodes structural relationships among adjacent camera views and transfers them to enhance cross-view depth consistency. Experiments on DDAD and nuScenes demonstrate the effectiveness of our method compared to conventional supervised methods and existing knowledge distillation approaches. Moreover, our method achieves a favorable trade-off between performance and efficiency, meeting real-time requirements.

单目深度知识蒸馏全景感知实时系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。