高效优化深度网络在多精度硬件上的混合精度部署
SEADA: An efficient methodology for optimizing mixed-precision DNNs on multi-precision spatial architectures

- 基于位级熵的逐层精度选择方法
- 可快速找到近优映射,降低延迟与能耗
- 适合硬件设计者进行多精度架构探索
混合精度计算已成为降低深度神经网络(DNN)延迟、能耗和内存占用的有效手段。然而,将混合精度网络高效映射到多精度空间架构面临诸多挑战:包括确定每层的合适精度、平衡各层对量化敏感度与架构异构性及系统约束之间的关系,以及准确估算异构精度分配下的系统级开销。本文提出SEADA,一种高效的方法论以应对上述问题。SEADA包含:(i) 可配置的多精度空间加速器系统级分析成本模型;(ii) 快速映射工具,用于在目标整数加速器上识别近优映射;(iii) 针对浮点层的分析模型,估算混合精度执行的整体收益;(iv) 基于位级熵的逐层精度选择方法,支持跨多种数值精度的高效分配。SEADA的高效性为多精度架构的设计空间探索提供了稳健框架。
原文摘要 · Abstract (English)
Mixed-precision computation has been introduced in deep neural networks (DNNs) as an effective approach to reduce latency, energy consumption, and memory footprint. However, efficiently mapping mixed-precision networks onto multi-precision spatial architectures poses several challenges. These include determining the appropriate precision for each layer, balancing layer-wise accuracy sensitivity to quantization against architectural heterogeneity and system-level constraints, and accurately estimating the system-level cost of heterogeneous precision assignments. This work presents SEADA, an efficient methodology designed to address these challenges. SEADA comprises: (i) a configurable system-level analytical cost model of a multi-precision spatial accelerator architecture; (ii) a fast mapping tool that identifies near-optimal mappings of DNN workloads onto the target integer accelerator; (iii) analytical models for floating-point layers to estimate the overall benefits of mixed-precision execution; and (iv) a per-layer precision selection methodology based on bit-level entropy, enabling efficient assignment across multiple numerical precisions. SEADA's efficiency provides designers with a robust framework for the design-space exploration of multi-precision architectures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。