通过鸟瞰图精简多模态输入,提升自动驾驶感知效率
Learning Content-Aware Multi-Modal Joint Input Pruning via Bird's-Eye-View Representation
- 用鸟瞰图统一多传感器数据,识别并剔除冗余输入区域
- 在NuScenes上实现计算量大幅降低,精度几乎不变
- 适合部署在算力受限的车载系统,尤其对实时性要求高场景
在自动驾驶领域,鸟瞰图(BEV)表示近年来受到广泛关注,成为多模态传感器融合的革新框架。该方法将传感器融合从规则驱动转向数据驱动,更有效地从异构传感器中提取特征。然而,基于BEV的技术通常带来显著的计算开销,需要高性能硬件支持,限制了实际应用。为此,本文提出一种内容感知的多模态联合输入剪枝方法。利用BEV作为共享参考,算法在感知模型主干网络前自动识别并移除非关键传感器区域。在NuScenes数据集上的大量实验验证了该方法的有效性,实现了显著的计算效率提升,且未牺牲感知精度。据我们所知,这是首个从输入剪枝角度缓解计算负担的工作。
原文摘要 · Abstract (English)
In the landscape of autonomous driving, Bird's-Eye-View (BEV) representation has recently garnered substantial academic attention, serving as a transformative framework for the fusion of multi-modal sensor inputs. This BEV paradigm effectively shifts the sensor fusion challenge from a rule-based methodology to a data-centric approach, thereby facilitating more nuanced feature extraction from an array of heterogeneous sensors. Notwithstanding its evident merits, the computational overhead associated with BEV-based techniques often mandates high-capacity hardware infrastructures, thus posing challenges for practical, real-world implementations. To mitigate this limitation, we introduce a novel content-aware multi-modal joint input pruning technique. Our method leverages BEV as a shared anchor to algorithmically identify and eliminate non-essential sensor regions prior to their introduction into the perception model's backbone. We validatethe efficacy of our approach through extensive experiments on the NuScenes dataset, demonstrating substantial computational efficiency without sacrificing perception accuracy. To the best of our knowledge, this work represents the first attempt to alleviate the computational burden from the input pruning point.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。