通过粒子化方法实现双因子稀疏性,提升低功耗DNN加速器的计算效率
BitParticle: Partializing Sparse Dual-Factors to Build Quasi-Synchronizing MAC Arrays for Energy-efficient DNNs
- 提出粒子化稀疏双因子处理机制,避免部分积爆炸
- 精确版本面积效率提升29.2%,近似版能效再增7.5%
- 引入准同步调度,减少流水线停顿,适合动态稀疏场景
量化深度神经网络(DNN)中的比特级稀疏性为优化乘累加(MAC)操作提供了巨大潜力。然而,两个关键挑战限制了其实际应用:一是传统比特串行方法无法同时利用两个因子的稀疏性,导致一个因子的稀疏性完全浪费;针对双因子稀疏性的方法仍处于早期探索阶段,面临部分积爆炸问题。二是比特级稀疏性的波动导致MAC操作周期数可变,现有同步调度方案灵活性差,仍造成大量MAC单元闲置。为此,本文提出一种基于粒子化的新式MAC单元,通过简单控制逻辑解决部分积爆炸问题,实现更省面积与能耗的设计。此外,舍弃次要中间结果可进一步简化硬件,仅带来轻微精度损失。为应对第二类挑战,引入准同步调度机制,在阵列中增加周期弹性,减少流水线停顿,提升MAC单元利用率。评估结果表明,所提精确版架构相比当前最优比特稀疏驱动架构,面积效率提升29.2%,能效相当;近似版相较精确版进一步提升能效7.5%。
原文摘要 · Abstract (English)
Bit-level sparsity in quantized deep neural networks (DNNs) offers significant potential for optimizing Multiply-Accumulate (MAC) operations. However, two key challenges still limit its practical exploitation. First, conventional bit-serial approaches cannot simultaneously leverage the sparsity of both factors, leading to a complete waste of one factor' s sparsity. Methods designed to exploit dual-factor sparsity are still in the early stages of exploration, facing the challenge of partial product explosion. Second, the fluctuation of bit-level sparsity leads to variable cycle counts for MAC operations. Existing synchronous scheduling schemes that are suitable for dual-factor sparsity exhibit poor flexibility and still result in significant underutilization of MAC units. To address the first challenge, this study proposes a MAC unit that leverages dual-factor sparsity through the emerging particlization-based approach. The proposed design addresses the issue of partial product explosion through simple control logic, resulting in a more area- and energy-efficient MAC unit. In addition, by discarding less significant intermediate results, the design allows for further hardware simplification at the cost of minor accuracy loss. To address the second challenge, a quasi-synchronous scheme is introduced that adds cycle-level elasticity to the MAC array, reducing pipeline stalls and thereby improving MAC unit utilization. Evaluation results show that the exact version of the proposed MAC array architecture achieves a 29.2% improvement in area efficiency compared to the state-of-the-art bit-sparsity-driven architecture, while maintaining comparable energy efficiency. The approximate variant further improves energy efficiency by 7.5%, compared to the exact version. Index-Terms: DNN acceleration, Bit-level sparsity, MAC unit
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。