arXiv:2409.07843cs.CVcs.RO2024-09被引 1

基于教师-学生学习的实时全景深度估计算法,适合边缘设备部署。

Real-time Multi-view Omnidirectional Depth Estimation for Real Scenarios based on Teacher-Student Learning with Unlabeled Data

  • 用教师模型生成真实数据伪标签,学生模型轻量化设计提升效率。
  • 在边缘设备上实现15帧/秒推理速度,精度接近顶尖方法。
  • 适用于自动驾驶、机器人等复杂室内外场景,泛化能力强。

全景深度估计可实现360度范围内的高效三维感知。然而,在自动驾驶和机器人等真实应用场景中,现有算法仍难以同时实现实时性能与跨场景鲁棒性。本文提出一种面向边缘计算平台的实时全景深度估计方法Rt-OmniMVS,引入联合球面扫描(Combined Spherical Sweeping)方法并设计轻量级网络结构,实现边缘设备上的实时运行。为提升精度、鲁棒性和泛化能力,采用教师-学生学习策略:以高精度立体匹配作为教师模型,为未标注的真实数据生成伪标签,并结合数据与模型增强技术训练学生模型。还构建了基于多视角鱼眼相机与边缘计算设备的HexaMODE全景深度感知系统。收集了一个包含真实世界与合成数据的大规模混合数据集用于训练。在公开数据集上的实验表明,所提方法性能接近当前最优,但资源消耗显著降低。系统在多种复杂真实场景(室内外)中均表现出高精度,边缘平台推理速度达15帧/秒。

原文摘要 · Abstract (English)

Omnidirectional depth estimation enables efficient 3D perception over a full 360-degree range. However, in real-world applications such as autonomous driving and robotics, achieving real-time performance and robust cross-scene generalization remains a significant challenge for existing algorithms. In this paper, we propose a real-time omnidirectional depth estimation method for edge computing platforms named Rt-OmniMVS, which introduces the Combined Spherical Sweeping method and implements the lightweight network structure to achieve real-time performance on edge computing platforms. To achieve high accuracy, robustness, and generalization in real-world environments, we introduce a teacher-student learning strategy. We leverage the high-precision stereo matching method as the teacher model to predict pseudo labels for unlabeled real-world data, and utilize data and model augmentation techniques for training to enhance performance of the student model Rt-OmniMVS. We also propose HexaMODE, an omnidirectional depth sensing system based on multi-view fisheye cameras and edge computation device. A large-scale hybrid dataset contains both unlabeled real-world data and synthetic data is collected for model training. Experiments on public datasets demonstrate that proposed method achieves results comparable to state-of-the-art approaches while consuming significantly less resource. The proposed system and algorithm also demonstrate high accuracy in various complex real-world scenarios, both indoors and outdoors, achieving an inference speed of 15 frames per second on edge computing platforms.

全景深度边缘计算教师学生

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。