提出统一设计空间,提升张量化模型在边缘设备的部署效率。
Comprehensive Design Space Exploration for Tensorized Neural Network Hardware Accelerators
- 联合优化计算路径、硬件架构与数据流映射,实现端到端高效部署。
- 在FPGA上实现,推理和训练延迟分别降低4倍和3.85倍。
- 适合关注边缘AI硬件加速的工程师与研究者。
高阶张量分解被广泛用于生成适用于边缘部署的紧凑型深度神经网络。然而,现有研究主要关注其算法优势,如精度和压缩率,而忽视了硬件部署效率。这类硬件无关的设计常掩盖张量化模型潜在的延迟与能耗优势。尽管部分工作通过优化收缩序列以减少乘加操作数来降低计算成本,但通常忽略底层硬件特性,导致实际性能不佳。我们观察到,收缩路径、硬件架构与数据流映射三者紧密耦合,必须在统一的设计空间内协同优化,才能最大化真实设备上的部署效率。为此,我们提出一种协同探索框架,将这些维度统一于一个设计空间中,面向边缘平台实现张量化神经网络的高效训练与推理。该框架定义以延迟为导向的搜索目标,并通过全局延迟驱动的探索实现端到端模型效率优化。优化后的配置在可配置FPGA核上实现,相比密集基线,推理和训练延迟分别降低最多4倍和3.85倍。
原文摘要 · Abstract (English)
High-order tensor decomposition has been widely adopted to obtain compact deep neural networks for edge deployment. However, existing studies focus primarily on its algorithmic advantages such as accuracy and compression ratio-while overlooking the hardware deployment efficiency. Such hardware-unaware designs often obscure the potential latency and energy benefits of tensorized models. Although several works attempt to reduce computational cost by optimizing the contraction sequence based on the number of multiply-accumulate operations, they typically neglect the underlying hardware characteristics, resulting in suboptimal real-world performance. We observe that the contraction path, hardware architecture, and dataflow mapping are tightly coupled and must be optimized jointly within a unified design space to maximize deployment efficiency on real devices. To this end, we propose a co-exploration framework that unifies these dimensions within a unified design space for efficient training and inference of tensorized neural networks on edge platforms. The framework formulates a latency oriented search objective and solves it via a global latency-driven exploration across the unified design space to achieve end-to-end model efficiency. The optimized configurations are implemented on a configurable FPGA kernel, achieving up to 4x and 3.85x lower inference and training latency compared with the dense baseline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。