提出资源感知的混合精度量化,让嵌入式FPGA上部署Transformer更灵活高效。
Resource-aware Mixed-precision Quantization for Enhancing Deployability of Transformers for Time-series Forecasting on Embedded FPGAs
- 通过可选存储资源类型提升VHDL模板灵活性,优化BRAM使用效率。
- 量化精度误差仅3%,实现接近实际部署的资源预估。
- 成功部署5个此前无法运行的混合精度模型,适合边缘AI开发者。
本研究针对在资源受限的嵌入式FPGA(Xilinx Spartan-7 XC7S15)上部署纯整数量化Transformer的挑战,通过引入可选中间结果存储资源类型,增强了VHDL模板的灵活性,从而高效利用BRAM,突破部署瓶颈。同时,提出一种资源感知的混合精度量化方法,使研究人员无需精通神经架构搜索即可探索硬件级量化策略。该方法对资源利用率的预测精度误差低至3%,与实际部署指标高度一致。相比先前工作,本方法成功实现了五种此前因统一位宽量化而无法部署的模型配置的部署,显著提升了Transformer在嵌入式系统中的可用性,推动了边缘设备上更多Transformer应用的落地。
原文摘要 · Abstract (English)
This study addresses the deployment challenges of integer-only quantized Transformers on resource-constrained embedded FPGAs (Xilinx Spartan-7 XC7S15). We enhanced the flexibility of our VHDL template by introducing a selectable resource type for storing intermediate results across model layers, thereby breaking the deployment bottleneck by utilizing BRAM efficiently. Moreover, we developed a resource-aware mixed-precision quantization approach that enables researchers to explore hardware-level quantization strategies without requiring extensive expertise in Neural Architecture Search. This method provides accurate resource utilization estimates with a precision discrepancy as low as 3%, compared to actual deployment metrics. Compared to previous work, our approach has successfully facilitated the deployment of model configurations utilizing mixed-precision quantization, thus overcoming the limitations inherent in five previously non-deployable configurations with uniform quantization bitwidths. Consequently, this research enhances the applicability of Transformers in embedded systems, facilitating a broader range of Transformer-powered applications on edge devices.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。