用贝塞尔曲线优化车道线检测,提升自动驾驶拓扑理解能力
TopoBDA: Towards Bezier Deformable Attention for Road Topology Understanding
- 引入贝塞尔控制点驱动的可变形注意力机制,精准捕捉细长车道线
- 在OpenLane-V2上实现顶尖车道中心线检测精度,3D车道检测也领先
- 适合做高精地图构建与自动驾驶感知系统的研发人员参考
道路拓扑理解对自动驾驶至关重要。本文提出TopoBDA(基于贝塞尔可变形注意力的道路拓扑理解),通过多摄像头360度图像生成鸟瞰图(BEV)特征,并利用采用贝塞尔可变形注意力(BDA)的Transformer解码器进行优化。BDA使用贝塞尔控制点驱动可变形注意力机制,显著提升对细长线状结构(如车道中心线)的检测与表示能力。此外,模型还引入实例掩码损失和一对多集合预测损失策略,进一步优化中心线检测与拓扑理解。在OpenLane-V2数据集上的实验表明,TopoBDA在车道中心线检测与拓扑推理任务中达到当前最优性能;在OpenLane-V1数据集上,3D车道检测也取得最佳结果。融合多模态数据(如LiDAR、雷达、SDMap)的实验显示,多源输入可进一步提升道路拓扑理解效果。
原文摘要 · Abstract (English)
Understanding road topology is crucial for autonomous driving. This paper introduces TopoBDA (Topology with Bezier Deformable Attention), a novel approach that enhances road topology comprehension by leveraging Bezier Deformable Attention (BDA). TopoBDA processes multi-camera 360-degree imagery to generate Bird's Eye View (BEV) features, which are refined through a transformer decoder employing BDA. BDA utilizes Bezier control points to drive the deformable attention mechanism, improving the detection and representation of elongated and thin polyline structures, such as lane centerlines. Additionally, TopoBDA integrates two auxiliary components: an instance mask formulation loss and a one-to-many set prediction loss strategy, to further refine centerline detection and enhance road topology understanding. Experimental evaluations on the OpenLane-V2 dataset demonstrate that TopoBDA outperforms existing methods, achieving state-of-the-art results in centerline detection and topology reasoning. TopoBDA also achieves the best results on the OpenLane-V1 dataset in 3D lane detection. Further experiments on integrating multi-modal data -- such as LiDAR, radar, and SDMap -- show that multimodal inputs can further enhance performance in road topology understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。