arXiv:2409.12008cs.CVcs.RO2024-09ICRA被引 1

同时预测未来场景的语义分割和深度图,提升机器人导航能力

Panoptic-Depth Forecasting

  • 用视觉图像联合预测未来帧的全景分割与深度图
  • 在KITTI-360和Cityscapes上实现优于基线的预测性能
  • 适合关注自动驾驶与机器人感知的读者

预测场景的语义和三维结构对机器人安全导航与行动规划至关重要。现有方法虽探索了语义与全景场景预测,但未考虑场景几何信息。本文提出全景-深度预测任务,从单目相机图像中联合预测未观测未来帧的全景分割与深度图。为支持该任务,我们扩展了流行的KITTI-360和Cityscapes数据集,利用激光雷达点云生成深度图,并使用序列标注数据。同时引入合适的评估指标,统一量化全景质量与深度估计精度。此外,我们提出两种基线并设计新型PDcast架构,通过基于Transformer的编码器、预测模块及任务专用解码器,学习丰富的时空表征以输出未来全景-深度结果。大量实验表明,PDcast在两个数据集和三个预测任务上均表现优异,有效应对主要挑战。代码已公开于https://pdcast.cs.uni-freiburg.de。

原文摘要 · Abstract (English)

Forecasting the semantics and 3D structure of scenes is essential for robots to navigate and plan actions safely. Recent methods have explored semantic and panoptic scene forecasting; however, they do not consider the geometry of the scene. In this work, we propose the panoptic-depth forecasting task for jointly predicting the panoptic segmentation and depth maps of unobserved future frames, from monocular camera images. To facilitate this work, we extend the popular KITTI-360 and Cityscapes benchmarks by computing depth maps from LiDAR point clouds and leveraging sequential labeled data. We also introduce a suitable evaluation metric that quantifies both the panoptic quality and depth estimation accuracy of forecasts in a coherent manner. Furthermore, we present two baselines and propose the novel PDcast architecture that learns rich spatio-temporal representations by incorporating a transformer-based encoder, a forecasting module, and task-specific decoders to predict future panoptic-depth outputs. Extensive evaluations demonstrate the effectiveness of PDcast across two datasets and three forecasting tasks, consistently addressing the primary challenges. We make the code publicly available at https://pdcast.cs.uni-freiburg.de.

全景分割深度预测视觉预测机器人感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。