自动化设计流程,让AI芯片在资源受限下高效运行。
MetaML-Pro: Cross-Stage Design Flow Automation for Efficient Deep Learning Acceleration
- 用元编程与贝叶斯优化打通上下层设计流
- 减少92% DSP和89% LUT使用,精度不变
- 适合想快速部署DNN加速器的工程师
本文提出一个统一框架,用于编码和自动化优化策略,以高效地将深度神经网络(DNN)部署在资源受限的硬件(如FPGA)上,同时保持高性能、高精度和资源效率。在该平台上部署DNN面临性能、资源占用(如DSPs和LUTs)和推理精度之间的平衡难题,通常需要大量人工干预和领域知识。本文新方法解决两个核心问题:(i) 编码自定义优化策略,(ii) 实现跨阶段优化搜索。框架将程序化DNN优化技术与基于高层次综合(HLS)的元编程结合,采用贝叶斯优化等先进设计空间探索(DSE)策略,自动化实现自顶向下和自底向上的设计流程,显著降低对人工干预和专业知识的依赖。此外,框架引入可定制的优化、变换和控制模块,提升DNN加速器的性能和资源效率。实验结果表明,对于部分网络,可实现高达92%的DSP和89%的LUT使用率降低,同时保持精度,优化时间相比网格搜索减少15.6倍。这些结果凸显了自动化生成资源高效DNN加速器设计的巨大潜力,且所需人力极少。
原文摘要 · Abstract (English)
This paper presents a unified framework for codifying and automating optimization strategies to efficiently deploy deep neural networks (DNNs) on resource-constrained hardware, such as FPGAs, while maintaining high performance, accuracy, and resource efficiency. Deploying DNNs on such platforms involves addressing the significant challenge of balancing performance, resource usage (e.g., DSPs and LUTs), and inference accuracy, which often requires extensive manual effort and domain expertise. Our novel approach addresses two core key issues: (i)~encoding custom optimization strategies and (ii)~enabling cross-stage optimization search. In particular, our proposed framework seamlessly integrates programmatic DNN optimization techniques with high-level synthesis (HLS)-based metaprogramming, leveraging advanced design space exploration (DSE) strategies like Bayesian optimization to automate both top-down and bottom-up design flows. Hence, we reduce the need for manual intervention and domain expertise. In addition, the framework introduces customizable optimization, transformation, and control blocks to enhance DNN accelerator performance and resource efficiency. Experimental results demonstrate up to a 92\% DSP and 89\% LUT usage reduction for select networks, while preserving accuracy, along with a 15.6-fold reduction in optimization time compared to grid search. These results highlight the potential for automating the generation of resource-efficient DNN accelerator designs with minimum effort.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。