构建首个针对FPGA加速器的资源与延迟预测基准,提升硬件设计迭代效率。
wa-hls4ml: A Benchmark and Surrogate Models for hls4ml Resource and Latency Estimation
- 基于hls4ml构建超68万神经网络的合成数据集,覆盖全连接与卷积结构
- 提出GNN与Transformer模型,对75%分位数的资源和延迟预测误差小于若干百分点
- 适合从事ML硬件加速、FPGA部署与自动化设计的研究者参考
随着机器学习在科学应用中日益依赖硬件实现以应对实时挑战,先进工具链的发展显著缩短了设计迭代时间。然而,此前未被重视的硬件综合过程现已成为快速迭代的新瓶颈。为此,我们提出wa-hls4ml,一个面向机器学习加速器资源与延迟估计的基准测试平台,及其配套的初始数据集,包含超过68万条使用hls4ml在Xilinx FPGA上合成的全连接与卷积神经网络。该基准评估了多种常见模型架构(主要来自科学领域)在资源与延迟预测上的表现,涵盖子集平均性能。同时,我们引入基于图神经网络(GNN)与Transformer的代理模型,用于预测加速器的资源占用与延迟。实验表明,这些模型在合成测试集上对75%分位数的预测结果与实际合成资源偏差在若干个百分点内。
原文摘要 · Abstract (English)
As machine learning (ML) is increasingly implemented in hardware to address real-time challenges in scientific applications, the development of advanced toolchains has significantly reduced the time required to iterate on various designs. These advancements have solved major obstacles, but also exposed new challenges. For example, processes that were not previously considered bottlenecks, such as hardware synthesis, are becoming limiting factors in the rapid iteration of designs. To mitigate these emerging constraints, multiple efforts have been undertaken to develop an ML-based surrogate model that estimates resource usage of ML accelerator architectures. We introduce wa-hls4ml, a benchmark for ML accelerator resource and latency estimation, and its corresponding initial dataset of over 680,000 fully connected and convolutional neural networks, all synthesized using hls4ml and targeting Xilinx FPGAs. The benchmark evaluates the performance of resource and latency predictors against several common ML model architectures, primarily originating from scientific domains, as exemplar models, and the average performance across a subset of the dataset. Additionally, we introduce GNN- and transformer-based surrogate models that predict latency and resources for ML accelerators. We present the architecture and performance of the models and find that the models generally predict latency and resources for the 75% percentile within several percent of the synthesized resources on the synthetic test dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。