arXiv:2412.11702cs.ARcs.CV2024-12被引 40

Flex-PE支持运行时可调精度与激活函数,兼顾能效与性能。

Flex-PE: Flexible and SIMD Multi-Precision Processing Element for AI Workloads

  • 设计可动态配置精度和激活函数的硬件单元,支持SIMD并行计算。
  • 在流水线模式下最高提升16倍吞吐,4位计算能效达8.42 GOPS/W。
  • 适合边缘与云端的AI推理、视觉变压器等高算力场景。

数据驱动型AI模型(如深度学习推理、训练、视觉变换器等)的快速发展,对支持运行时可配置精度与非线性激活函数的硬件提出了迫切需求。现有方案虽支持多样精度或激活函数重配置,但无法同时满足二者。本文提出一种灵活且支持SIMD的多精度处理单元(FlexPE),可动态支持sigmoid、tanh、ReLU、softmax等激活函数及乘加运算。该设计在流水线模式下实现最高16倍(FxP4)、8倍(FxP8)、4倍(FxP16)和1倍(FxP32)吞吐提升,且100%时间复用硬件资源。针对边缘端应用,提出面积高效的多精度迭代模式。在VGG16上,输入特征图和权重滤波器的DMA读取分别减少62倍和371倍,精度损失仅2%,能效达8.42 GOPS/W。架构支持新兴4位计算用于深度学习推理,同时提升FxP8/16模式在变换器等高性能计算中的吞吐表现,为未来边缘与云环境下的能效型AI加速器提供可行路径。

原文摘要 · Abstract (English)

The rapid adaptation of data driven AI models, such as deep learning inference, training, Vision Transformers (ViTs), and other HPC applications, drives a strong need for runtime precision configurable different non linear activation functions (AF) hardware support. Existing solutions support diverse precision or runtime AF reconfigurability but fail to address both simultaneously. This work proposes a flexible and SIMD multiprecision processing element (FlexPE), which supports diverse runtime configurable AFs, including sigmoid, tanh, ReLU and softmax, and MAC operation. The proposed design achieves an improved throughput of up to 16X FxP4, 8X FxP8, 4X FxP16 and 1X FxP32 in pipeline mode with 100% time multiplexed hardware. This work proposes an area efficient multiprecision iterative mode in the SIMD systolic arrays for edge AI use cases. The design delivers superior performance with up to 62X and 371X reductions in DMA reads for input feature maps and weight filters in VGG16, with an energy efficiency of 8.42 GOPS / W within the accuracy loss of 2%. The proposed architecture supports emerging 4-bit computations for DL inference while enhancing throughput in FxP8/16 modes for transformers and other HPC applications. The proposed approach enables future energy-efficient AI accelerators in edge and cloud environments.

AI加速器多精度计算边缘计算能效优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。