提出单阶段融合的雷达相机深度估计模型,提升效率与精度。
TacoDepth: Towards Efficient Radar-Camera Depth Estimation with One-stage Fusion
- 单阶段融合雷达点云图结构与图像特征,避免中间深度估计
- 相比最优方法,精度提升12.8%,推理速度提高91.8%
- 适用于自动驾驶等对实时性要求高的场景
雷达-相机深度估计旨在通过融合图像与雷达数据预测稠密且精确的度量深度。模型效率对自动驾驶和机器人平台实现实时处理至关重要。然而,由于雷达回波稀疏,现有方法多采用包含中间伪稠密深度的多阶段框架,导致计算耗时且鲁棒性差。为此,我们提出TacoDepth,一种高效且准确的单阶段雷达-相机深度估计模型。设计了基于图结构的雷达特征提取器和金字塔式雷达融合模块,有效捕捉并整合雷达点云的图结构信息,无需依赖中间深度结果,显著提升模型效率与鲁棒性。此外,TacoDepth支持多种推理模式,在速度与精度间实现更好平衡。大量实验验证了方法有效性:相比先前最先进方法,精度提升12.8%,处理速度提高91.8%。本工作为高效雷达-相机深度估计提供了新思路。
原文摘要 · Abstract (English)
Radar-Camera depth estimation aims to predict dense and accurate metric depth by fusing input images and Radar data. Model efficiency is crucial for this task in pursuit of real-time processing on autonomous vehicles and robotic platforms. However, due to the sparsity of Radar returns, the prevailing methods adopt multi-stage frameworks with intermediate quasi-dense depth, which are time-consuming and not robust. To address these challenges, we propose TacoDepth, an efficient and accurate Radar-Camera depth estimation model with one-stage fusion. Specifically, the graph-based Radar structure extractor and the pyramid-based Radar fusion module are designed to capture and integrate the graph structures of Radar point clouds, delivering superior model efficiency and robustness without relying on the intermediate depth results. Moreover, TacoDepth can be flexible for different inference modes, providing a better balance of speed and accuracy. Extensive experiments are conducted to demonstrate the efficacy of our method. Compared with the previous state-of-the-art approach, TacoDepth improves depth accuracy and processing speed by 12.8% and 91.8%. Our work provides a new perspective on efficient Radar-Camera depth estimation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。