自动构建快速精准的DNN加速器性能模型,提升边缘设备部署效率
Automatic Generation of Fast and Accurate Performance Models for Deep Neural Network Accelerators
- 基于依赖图分析,仅需154次循环即可估算41.9亿指令延迟
- 相比仿真,精度高于回归与解析模型,错误率更低
- 适合芯片架构设计者和AI推理优化工程师使用
在资源受限的边缘设备上部署深度神经网络(DNN)是一项挑战,需要定制化硬件加速器并清晰理解其执行特定人工智能工作负载时的性能特征。为此,我们提出一种自动化生成方法,可快速准确地估计DNN映射到系统建模、简洁描述的加速器架构上的延迟。通过我们的加速器描述方法,我们建模了代表性DNN加速器如Gemmini、UltraTrail、Plasticine衍生架构及可参数化阵列。结合这些架构的DNN映射,我们进行联合的DNN/硬件依赖图分析,最理想情况下仅需评估154个循环内核迭代,即可估算出41.9亿条指令的性能,实现显著加速。相较于仿真结果,该方法在平均绝对百分比误差(MAPE)上优于回归与解析模型,同时比RTL仿真快几个数量级。
原文摘要 · Abstract (English)
Implementing Deep Neural Networks (DNNs) on resource-constrained edge devices is a challenging task that requires tailored hardware accelerator architectures and a clear understanding of their performance characteristics when executing the intended AI workload. To facilitate this, we present an automated generation approach for fast performance models to accurately estimate the latency of a DNN mapped onto systematically modeled and concisely described accelerator architectures. Using our accelerator architecture description method, we modeled representative DNN accelerators such as Gemmini, UltraTrail, Plasticine-derived, and a parameterizable systolic array. Together with DNN mappings for those modeled architectures, we perform a combined DNN/hardware dependency graph analysis, which enables us, in the best case, to evaluate only 154 loop kernel iterations to estimate the performance for 4.19 billion instructions achieving a significant speedup. We outperform regression and analytical models in terms of mean absolute percentage error (MAPE) compared to simulation results, while being several magnitudes faster than an RTL simulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。