用稀疏高斯过程加速神经网络硬件联合搜索,提升设计效率与精度。
Coflex: Enhancing HW-NAS with Sparse Gaussian Processes for Efficient and Scalable DNN Accelerator Design
- 引入稀疏诱导点降低高斯过程复杂度,实现近线性计算开销。
- 在多个基准上达到更高网络精度和更低能时延积,提速1.9至9.5倍。
- 适合边缘AI芯片设计者快速探索高性能低功耗加速器架构。
硬件感知神经架构搜索(HW-NAS)是一种高效方法,可自动协同优化神经网络性能与硬件能效,特别适用于边缘端深度神经网络加速器的开发。然而,庞大的搜索空间和高昂的计算成本严重制约其实际应用。为此,我们提出Coflex,一种将稀疏高斯过程(SGP)与多目标贝叶斯优化相结合的新框架。通过使用稀疏诱导点,Coflex将高斯过程核函数复杂度从立方级降至近线性,且不牺牲优化性能。该方法实现了对大规模搜索空间的可扩展逼近,在显著降低计算开销的同时保持高预测精度。我们在多个基准上评估了Coflex的有效性,聚焦于面向加速器的架构设计。实验结果表明,Coflex在模型精度和能时延积方面优于现有最先进方法,并实现了1.9到9.5倍的计算速度提升。
原文摘要 · Abstract (English)
Hardware-Aware Neural Architecture Search (HW-NAS) is an efficient approach to automatically co-optimizing neural network performance and hardware energy efficiency, making it particularly useful for the development of Deep Neural Network accelerators on the edge. However, the extensive search space and high computational cost pose significant challenges to its practical adoption. To address these limitations, we propose Coflex, a novel HW-NAS framework that integrates the Sparse Gaussian Process (SGP) with multi-objective Bayesian optimization. By leveraging sparse inducing points, Coflex reduces the GP kernel complexity from cubic to near-linear with respect to the number of training samples, without compromising optimization performance. This enables scalable approximation of large-scale search space, substantially decreasing computational overhead while preserving high predictive accuracy. We evaluate the efficacy of Coflex across various benchmarks, focusing on accelerator-specific architecture. Our experimental results show that Coflex outperforms state-of-the-art methods in terms of network accuracy and Energy-Delay-Product, while achieving a computational speed-up ranging from 1.9x to 9.5x.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。