用芯片组设计提升自动驾驶感知算力,效率更高。
Performance Implications of Multi-Chiplet Neural Processing Units on Autonomous Driving Perception
- 拆解感知模型,在多芯片组加速器上调度部署。
- 相比单体芯片,吞吐量提升82%,引擎利用率提高2.8倍。
- 适合车载AI算力优化,对芯片设计者有参考价值。
我们研究了新兴的基于芯片组的神经网络处理单元在受限车载环境中的应用,以加速车辆AI感知任务。芯片组技术因其在性能、模块化和定制化间的成本效益平衡,正成为车载架构的关键;而感知模型是自动驾驶系统中最耗算力的任务。以特斯拉自动辅助驾驶感知流程为案例,我们分解其组件模型,并在不同芯片组加速器上进行性能分析。基于这些洞察,提出一种新型调度策略,高效部署感知任务于多芯片组AI加速器。使用标准DNN性能模拟器MAESTRO的实验表明,相比单体加速器设计,该方法实现82%的吞吐量提升和2.8倍的处理引擎利用率提升。
原文摘要 · Abstract (English)
We study the application of emerging chiplet-based Neural Processing Units to accelerate vehicular AI perception workloads in constrained automotive settings. The motivation stems from how chiplets technology is becoming integral to emerging vehicular architectures, providing a cost-effective trade-off between performance, modularity, and customization; and from perception models being the most computationally demanding workloads in a autonomous driving system. Using the Tesla Autopilot perception pipeline as a case study, we first breakdown its constituent models and profile their performance on different chiplet accelerators. From the insights, we propose a novel scheduling strategy to efficiently deploy perception workloads on multi-chip AI accelerators. Our experiments using a standard DNN performance simulator, MAESTRO, show our approach realizes 82% and 2.8x increase in throughput and processing engines utilization compared to monolithic accelerator designs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。