为混合模型设计高效部署方案,提升AR/VR设备的响应速度与能效。
Neural Architecture Search of Hybrid Models for NPU-CIM Heterogeneous AR/VR Devices
- 利用NPU与存算一体芯片的异构特性,优化混合模型执行
- 在ImageNet上实现最高1.34%的精度提升
- 适合开发低延迟、低功耗边缘AI系统的工程师
低延迟、低功耗边缘AI对虚拟现实与增强现实应用至关重要。近期研究表明,结合卷积层(CNN)与视觉变压器(ViT)的混合模型在多种计算机视觉与机器学习任务中常能取得更优的准确率/性能权衡。然而,由于数据流和内存访问模式差异大,这类混合模型在系统层面可能带来延迟与能效挑战。本文利用神经处理单元(NPU)与存算一体(CIM)的架构异构性,采用多样化执行策略以高效运行混合模型。提出H4H-NAS框架,用于为兼具NPU与CIM的异构边缘系统设计高效混合CNN/ViT模型。该框架基于真实硅片测量的NPU性能与行业IP估算的CIM性能构建性能评估器,实现细粒度搜索,在ImageNet上最高实现1.34%的top-1精度提升。算法与硬件协同设计结果显示,相比基线方案,整体延迟降低56.08%,能耗减少41.72%。该框架指导了混合网络结构与异构系统架构的设计。
原文摘要 · Abstract (English)
Low-Latency and Low-Power Edge AI is essential for Virtual Reality and Augmented Reality applications. Recent advances show that hybrid models, combining convolution layers (CNN) and transformers (ViT), often achieve superior accuracy/performance tradeoff on various computer vision and machine learning (ML) tasks. However, hybrid ML models can pose system challenges for latency and energy-efficiency due to their diverse nature in dataflow and memory access patterns. In this work, we leverage the architecture heterogeneity from Neural Processing Units (NPU) and Compute-In-Memory (CIM) and perform diverse execution schemas to efficiently execute these hybrid models. We also introduce H4H-NAS, a Neural Architecture Search framework to design efficient hybrid CNN/ViT models for heterogeneous edge systems with both NPU and CIM. Our H4H-NAS approach is powered by a performance estimator built with NPU performance results measured on real silicon, and CIM performance based on industry IPs. H4H-NAS searches hybrid CNN/ViT models with fine granularity and achieves significant (up to 1.34%) top-1 accuracy improvement on ImageNet dataset. Moreover, results from our Algo/HW co-design reveal up to 56.08% overall latency and 41.72% energy improvements by introducing such heterogeneous computing over baseline solutions. The framework guides the design of hybrid network architectures and system architectures of NPU+CIM heterogeneous systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。