提出高效硬件感知的神经网络架构搜索代理模型框架
ESM: A Framework for Building Effective Surrogate Models for Hardware-Aware Neural Architecture Search
- 构建面向GPU设备的延迟预测代理模型框架
- 系统分析影响预测精度的关键因素并优化全流程成本
- 适合做边缘设备模型优化的研究者与工程师参考
硬件感知神经架构搜索(NAS)是为资源受限设备设计高效深度神经网络的前沿技术。代理模型在其中起关键作用,可快速预测候选模型在目标硬件上的性能指标(如推理延迟和能耗)。本文聚焦于构建硬件感知的延迟预测模型,系统研究了不同类型的代理模型及其优劣,并深入分析影响预测准确性的各类因素,旨在评估模型设计各阶段的重要性,识别出适用于GPU设备的有效建模方法与训练策略。基于上述洞察,提出一个全流程框架,兼顾数据生成与模型构建的可靠性与效率,综合考虑各阶段的整体开销。
原文摘要 · Abstract (English)
Hardware-aware Neural Architecture Search (NAS) is one of the most promising techniques for designing efficient Deep Neural Networks (DNNs) for resource-constrained devices. Surrogate models play a crucial role in hardware-aware NAS as they enable efficient prediction of performance characteristics (e.g., inference latency and energy consumption) of different candidate models on the target hardware device. In this paper, we focus on building hardware-aware latency prediction models. We study different types of surrogate models and highlight their strengths and weaknesses. We perform a systematic analysis to understand the impact of different factors that can influence the prediction accuracy of these models, aiming to assess the importance of each stage involved in the model designing process and identify methods and policies necessary for designing/training an effective estimation model, specifically for GPU-powered devices. Based on the insights gained from the analysis, we present a holistic framework that enables reliable dataset generation and efficient model generation, considering the overall costs of different stages of the model generation pipeline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。