arXiv:2411.11466cs.CV2024-11

统一单目几何理解,实时分割与深度估计一次完成。

MGNiceNet: Unified Monocular Geometric Scene Understanding

  • 用联动卷积核统一建模分割与深度预测
  • 在Cityscapes和KITTI上达实时方法最优性能
  • 无需视频分割标注,适合自动驾驶部署

单目几何场景理解结合全景分割与自监督深度估计,聚焦自动驾驶中的实时应用。本文提出MGNiceNet,一种基于先进实时全景分割方法RT-K-Net的统一框架,扩展其结构以同时实现全景分割与自监督单目深度估计。通过引入紧密耦合的自监督深度预测模块,显式利用全景路径信息进行深度推断,并设计了全景引导的运动掩码方法,在不依赖视频全景分割标注的前提下提升深度估计效果。在Cityscapes和KITTI两个主流自动驾驶数据集上评估,MGNiceNet在实时方法中表现领先,显著缩小与计算开销更大的方法之间的差距。源代码与训练模型已开源。

原文摘要 · Abstract (English)

Monocular geometric scene understanding combines panoptic segmentation and self-supervised depth estimation, focusing on real-time application in autonomous vehicles. We introduce MGNiceNet, a unified approach that uses a linked kernel formulation for panoptic segmentation and self-supervised depth estimation. MGNiceNet is based on the state-of-the-art real-time panoptic segmentation method RT-K-Net and extends the architecture to cover both panoptic segmentation and self-supervised monocular depth estimation. To this end, we introduce a tightly coupled self-supervised depth estimation predictor that explicitly uses information from the panoptic path for depth prediction. Furthermore, we introduce a panoptic-guided motion masking method to improve depth estimation without relying on video panoptic segmentation annotations. We evaluate our method on two popular autonomous driving datasets, Cityscapes and KITTI. Our model shows state-of-the-art results compared to other real-time methods and closes the gap to computationally more demanding methods. Source code and trained models are available at https://github.com/markusschoen/MGNiceNet.

单目深度全景分割自动驾驶实时推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。