用可配置模块和多项式近似,让FPGA高效部署卷积网络
Implémentation Efficiente de Fonctions de Convolution sur FPGA à l'Aide de Blocs Paramétrables et d'Approximations Polynomiales
- 设计可调参数的卷积硬件模块,适配不同资源约束
- 建立数学模型精准预测FPGA资源占用,误差小
- 适合需要低延迟、高能效的嵌入式深度学习部署
在现场可编程门阵列(FPGA)上实现卷积神经网络(CNN)已成为替代GPU的有前景方案,具备更低延迟、更高能效和更强灵活性。然而,由于所需硬件知识复杂以及综合、布局布线周期长,设计迭代困难,严重限制了网络结构的快速探索,尤其在资源极度受限条件下难以优化。本文提出一套可配置卷积模块库,用于优化FPGA实现并适应可用资源;同时构建方法论框架,建立可预测FPGA资源消耗的数学模型。通过分析参数相关性与误差指标验证,结果表明所设计模块能有效适配硬件约束,且模型对资源消耗预测准确,为FPGA选型与优化部署提供有力工具。
原文摘要 · Abstract (English)
Implementing convolutional neural networks (CNNs) on field-programmable gate arrays (FPGAs) has emerged as a promising alternative to GPUs, offering lower latency, greater power efficiency and greater flexibility. However, this development remains complex due to the hardware knowledge required and the long synthesis, placement and routing stages, which slow down design cycles and prevent rapid exploration of network configurations, making resource optimisation under severe constraints particularly challenging. This paper proposes a library of configurable convolution Blocks designed to optimize FPGA implementation and adapt to available resources. It also presents a methodological framework for developing mathematical models that predict FPGA resources utilization. The approach is validated by analyzing the correlation between the parameters, followed by error metrics. The results show that the designed blocks enable adaptation of convolution layers to hardware constraints, and that the models accurately predict resource consumption, providing a useful tool for FPGA selection and optimized CNN deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。