用CORDIC实现高效可重构的AI加速,性能提升4.6倍
CORDIC Is All You Need
- 采用流水线CORDIC模块处理线性计算与非线性激活函数
- 在28nm下吞吐量提升4.64倍,功耗和面积分别降低5.02倍和4.06倍
- 适用于Transformer、RNN、DNN等场景,适合边缘AI部署
人工智能需要高效的可适配硬件加速器以支持高吞吐量百万级运算。本文提出基于可重构处理引擎(RPE)的流水线架构,集成CORDIC模块用于线性乘累加(MAC)计算及非线性迭代激活函数(如tanh、sigmoid、softmax)。该架构在40%剪枝率下,实现高达4.64倍的吞吐量提升,功耗与面积分别降低5.02倍和4.06倍(CMOS 28 nm),精度损失微小。FPGA实现显示资源节省最高达2.5倍,功耗降低3倍。所提出的可重构增强型流式CORDIC引擎(SYCore)采用输出驻留数据流与CAESAR控制引擎,支持Transformer、RNN/LSTM、DNN等多样化AI任务,适用于图像检测、大语言模型和语音识别等应用。该节能灵活方案扩展了对新兴工作负载的支持,适用于边缘AI加速器。
原文摘要 · Abstract (English)
Artificial intelligence necessitates adaptable hardware accelerators for efficient high-throughput million operations. We present pipelined architecture with CORDIC block for linear MAC computations and nonlinear iterative Activation Functions (AF) such as $tanh$, $sigmoid$, and $softmax$. This approach focuses on a Reconfigurable Processing Engine (RPE) based systolic array, with 40\% pruning rate, enhanced throughput up to 4.64$\times$, and reduction in power and area by 5.02 $\times$ and 4.06 $\times$ at CMOS 28 nm, with minor accuracy loss. FPGA implementation achieves a reduction of up to 2.5 $\times$ resource savings and 3 $\times$ power compared to prior works. The Systolic CORDIC engine for Reconfigurability and Enhanced throughput (SYCore) deploys an output stationary dataflow with the CAESAR control engine for diverse AI workloads such as Transformers, RNNs/LSTMs, and DNNs for applications like image detection, LLMs, and speech recognition. The energy-efficient and flexible approach extends the enhanced approach for edge AI accelerators supporting emerging workloads.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。