用自监督深度估计实现无需标定的单目3D车道检测
Depth3DLane: Fusing Monocular 3D Lane Detection with Self-Supervised Monocular Depth Estimation
- 双路径结构融合前视语义与鸟瞰空间信息
- 在OpenLane上达到竞品水平,且不依赖相机参数
- 适合无标定场景,如众包高精地图构建
单目3D车道检测对自动驾驶至关重要,但因缺乏显式空间信息而困难。现有方法依赖昂贵深度传感器或需真实深度标签,难以规模化;且通常假设相机参数已知,限制了在众包高清地图等场景的应用。为此,我们提出Depth3DLane,一种新颖的双路径框架,通过自监督单目深度估计获得场景点云表示,无需额外传感器或真值深度数据。其鸟瞰路径提取显式空间信息,前视路径同步捕获丰富语义。模型利用3D车道锚框从两路径采样特征并推断精确3D车道几何。此外,框架可逐帧预测相机参数,并引入理论驱动的拟合流程提升分段稳定性。大量实验表明,Depth3DLane在OpenLane基准上表现优异;使用学习参数替代真值参数后,可在无标定条件下应用,优于以往方法。
原文摘要 · Abstract (English)
Monocular 3D lane detection is essential for autonomous driving, but challenging due to the inherent lack of explicit spatial information. Multi-modal approaches rely on expensive depth sensors, while methods incorporating fully-supervised depth networks rely on ground-truth depth data that is impractical to collect at scale. Additionally, existing methods assume that camera parameters are available, limiting their applicability in scenarios like crowdsourced high-definition (HD) lane mapping. To address these limitations, we propose Depth3DLane, a novel dual-pathway framework that integrates self-supervised monocular depth estimation to provide explicit structural information, without the need for expensive sensors or additional ground-truth depth data. Leveraging a self-supervised depth network to obtain a point cloud representation of the scene, our bird's-eye view pathway extracts explicit spatial information, while our front view pathway simultaneously extracts rich semantic information. Depth3DLane then uses 3D lane anchors to sample features from both pathways and infer accurate 3D lane geometry. Furthermore, we extend the framework to predict camera parameters on a per-frame basis and introduce a theoretically motivated fitting procedure to enhance stability on a per-segment basis. Extensive experiments demonstrate that Depth3DLane achieves competitive performance on the OpenLane benchmark dataset. Furthermore, experimental results show that using learned parameters instead of ground-truth parameters allows Depth3DLane to be applied in scenarios where camera calibration is infeasible, unlike previous methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。