提出全稀疏框架SparseAlign,提升协同感知效率与精度。
SparseAlign: A Fully Sparse Framework for Cooperative Object Detection
- 采用稀疏3D主干+查询式时序建模,降低计算开销。
- 在OPV2V/DairV2X上超越现有方法,通信量更低。
- 适合长距离协同检测,对稀疏特征优化更鲁棒。
协同感知可扩大自车视域、减少遮挡,从而提升自动驾驶的感知性能与安全性。尽管已有研究在协同目标检测上取得成功,但多数方法基于密集的鸟瞰图(BEV)特征图,计算开销大,难以扩展至长距离检测问题。高效全稀疏框架仍较少被探索。本文设计全稀疏框架SparseAlign,包含三个关键组件:增强型稀疏3D主干网络、基于查询的时序上下文学习模块,以及专为稀疏特征设计的鲁棒检测头。在OPV2V和DairV2X数据集上的大量实验表明,该框架虽具稀疏性,但仍优于当前最优方法,且通信带宽需求更低。此外,在面向时间对齐的协同检测任务中,OPV2Vt和DairV2Xt数据集上的实验也显示显著性能提升。
原文摘要 · Abstract (English)
Cooperative perception can increase the view field and decrease the occlusion of an ego vehicle, hence improving the perception performance and safety of autonomous driving. Despite the success of previous works on cooperative object detection, they mostly operate on dense Bird's Eye View (BEV) feature maps, which are computationally demanding and can hardly be extended to long-range detection problems. More efficient fully sparse frameworks are rarely explored. In this work, we design a fully sparse framework, SparseAlign, with three key features: an enhanced sparse 3D backbone, a query-based temporal context learning module, and a robust detection head specially tailored for sparse features. Extensive experimental results on both OPV2V and DairV2X datasets show that our framework, despite its sparsity, outperforms the state of the art with less communication bandwidth requirements. In addition, experiments on the OPV2Vt and DairV2Xt datasets for time-aligned cooperative object detection also show a significant performance gain compared to the baseline works.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。