arXiv:2410.06516cs.ROcs.AI2024-10被引 2

四任务合一的高效车载感知框架,提升实时性与资源利用率。

QuadBEV: An Efficient Quadruple-Task Perception Framework via Bird's-Eye-View Representation

  • 共享主干网络,统一处理四类感知任务
  • 减少冗余计算,适配嵌入式系统部署
  • 解决多任务学习冲突,提升模型稳定性

鸟瞰图(BEV)感知因能整合多传感器输入并生成统一表征,已成为自动驾驶系统的关键组件,显著提升下游任务性能。然而,现有BEV模型计算开销大,难以在资源受限的车载设备中部署。为此,本文提出QuadBEV,一种高效的多任务感知框架,通过共享空间与上下文信息,联合处理四个核心任务:3D目标检测、车道线检测、地图分割和占据预测。该框架采用共享主干网络与任务专用头结构,有效减少重复计算,并缓解多任务学习中的学习率敏感性与任务目标冲突问题。实验表明,QuadBEV在保持高性能的同时显著提升系统效率,具备良好的鲁棒性与实际应用潜力。

原文摘要 · Abstract (English)

Bird's-Eye-View (BEV) perception has become a vital component of autonomous driving systems due to its ability to integrate multiple sensor inputs into a unified representation, enhancing performance in various downstream tasks. However, the computational demands of BEV models pose challenges for real-world deployment in vehicles with limited resources. To address these limitations, we propose QuadBEV, an efficient multitask perception framework that leverages the shared spatial and contextual information across four key tasks: 3D object detection, lane detection, map segmentation, and occupancy prediction. QuadBEV not only streamlines the integration of these tasks using a shared backbone and task-specific heads but also addresses common multitask learning challenges such as learning rate sensitivity and conflicting task objectives. Our framework reduces redundant computations, thereby enhancing system efficiency, making it particularly suited for embedded systems. We present comprehensive experiments that validate the effectiveness and robustness of QuadBEV, demonstrating its suitability for real-world applications.

BEV感知多任务学习自动驾驶嵌入式部署

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。