针对边缘设备动态推理,提出软硬件协同优化框架,显著降低能耗延迟积。
Hardware-Algorithm Co-Optimization of Early-Exit Neural Networks for Multi-Core Edge Accelerators
- 构建软硬件协同设计框架,联合优化退出位置、量化精度与多核映射
- 在8位量化下,能耗延迟积比静态基线降低超50%
- 适合追求边缘端高效动态推理的系统设计者与模型部署工程师
将动态神经网络部署于边缘加速器时,需考虑超越传统计算量(如乘加操作)的硬件约束。在早期退出神经网络(EENN)中,退出位置、量化级别与硬件工作负载映射之间存在复杂非线性交互,影响内存流量、加速器利用率,最终决定能效-延迟权衡。现有神经架构搜索(NAS)方法通常依赖代理指标或硬件闭环评估,难以充分理解这些交互。本文提出一种面向EENN的软硬件协同设计框架,显式建模量化、退出配置与多核加速器映射之间的相互作用。通过分析性设计空间探索,揭示微小架构变化因张量维度对齐与数据流效应可引发显著硬件效率波动。基于此分析,将EENN部署建模为兼顾准确率、能效-延迟积、退出开销与动态推理行为的约束多目标优化问题。在CIFAR-10上的实验表明,所提框架在8位量化下实现的能效-延迟积相比静态基线降低超过50%。结果凸显了在异构边缘平台进行部署感知协同设计的重要性。
原文摘要 · Abstract (English)
Deployment of dynamic neural networks on edge accelerators requires careful consideration of hardware constraints beyond conventional complexity metrics such as Multiply-Accumulate operations. In Early-Exiting Neural Networks (EENN), exit placement, quantization level, and hardware workload mapping interact in non-trivial ways, influencing memory traffic, accelerator utilization, and ultimately energy-latency trade-offs. These interactions remain insufficiently understood in existing Neural Architecture Search (NAS) approaches, which typically rely on proxy metrics or hardware-in-the-loop evaluation. This work presents a hardware-algorithm co-design framework for EENN that explicitly models the interplay between quantization, exit configuration, and multi-core accelerator mapping. Using analytical design space exploration, we characterize how small architectural variations can induce disproportionate changes in hardware efficiency due to tensor dimension alignment and dataflow effects. Building on this analysis, we formulate EENN deployment as a constrained multi-objective optimization problem balancing accuracy, energy-latency product, exit overhead, and dynamic inference behavior. Experimental results on CIFAR-10 demonstrate that the proposed framework identifies architectures achieving over 50\% reduction in energy-latency product compared to static baselines under 8-bit quantization. The results highlight the importance of deployment-aware co-design for dynamic inference on heterogeneous edge platforms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。