SNAC-Pack让神经网络自动适配FPGA,兼顾精度与资源效率。
Surrogate Neural Architecture Codesign Package (SNAC-Pack)
- 结合多阶段搜索与资源延迟估计算法,实现精度、资源、延迟三重优化。
- 在高能物理任务中达63.84%准确率,部署后资源利用与传统方法相当。
- 适合需要低功耗、高效率的FPGA嵌入式部署场景,开源可用。
神经架构搜索可自动化模型设计,但现有方法难以精准优化真实硬件性能,常依赖如浮点运算量等代理指标。本文提出集成框架Surrogate Neural Architecture Codesign Package(SNAC-Pack),自动发现并优化适用于FPGA部署的神经网络。SNAC-Pack融合神经架构协同设计的多阶段搜索能力与资源利用率及延迟估计算法,实现在不进行每次候选模型耗时综合的情况下,对精度、FPGA资源利用率和延迟进行多目标优化。我们在高能物理喷注分类任务上验证了SNAC-Pack,实现63.84%的准确率。在Xilinx Virtex UltraScale+ VU13P FPGA上综合后,其性能与基线模型相当,且资源利用率与采用传统BOPs指标优化的模型相近。该工作展示了硬件感知神经架构搜索在资源受限部署中的潜力,并提供了开源框架以自动化高效FPGA加速模型的设计。
原文摘要 · Abstract (English)
Neural Architecture Search is a powerful approach for automating model design, but existing methods struggle to accurately optimize for real hardware performance, often relying on proxy metrics such as bit operations. We present Surrogate Neural Architecture Codesign Package (SNAC-Pack), an integrated framework that automates the discovery and optimization of neural networks focusing on FPGA deployment. SNAC-Pack combines Neural Architecture Codesign's multi-stage search capabilities with the Resource Utilization and Latency Estimator, enabling multi-objective optimization across accuracy, FPGA resource utilization, and latency without requiring time-intensive synthesis for each candidate model. We demonstrate SNAC-Pack on a high energy physics jet classification task, achieving 63.84% accuracy with resource estimation. When synthesized on a Xilinx Virtex UltraScale+ VU13P FPGA, the SNAC-Pack model matches baseline accuracy while maintaining comparable resource utilization to models optimized using traditional BOPs metrics. This work demonstrates the potential of hardware-aware neural architecture search for resource-constrained deployments and provides an open-source framework for automating the design of efficient FPGA-accelerated models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。