优化可微逻辑门网络在FPGA上的资源利用与速度平衡
Resource Utilization of Differentiable Logic Gate Networks Deployed on FPGAs

- 通过调整网络深度和宽度,研究FPGA上LGN的硬件特性
- 末层变窄可降低28%资源占用并提升时序性能
- 为FPGA部署提供可量化的架构选择依据
边缘机器学习常追求小模型的高智能,同时最小化电路尺寸和功耗。可微逻辑门网络(LGN)在保持纳秒级预测速度的同时,相比传统二值神经网络显著减少了资源需求。然而,LGN参数与FPGA硬件综合结果之间的权衡尚未明确。本文研究了在FPGA上合成LGN时,深度与宽度变化对功耗、资源利用率、推理速度和模型精度的影响。结果表明,末层对时序和资源使用至关重要,其变窄可使资源用量减少28%,因该层决定求和操作的逻辑规模。在时序和布线约束下,当末层较窄时,可合成更深更宽的LGN。本文进一步给出了基于可用查找表(LUT)数量的架构选型权衡建议,助力ML工程师高效设计。
原文摘要 · Abstract (English)
On-edge machine learning (ML) often strives to maximize the intelligence of small models while miniaturizing the circuit size and power needed to perform inference. Meeting these needs, differentiable Logic Gate Networks (LGN) have demonstrated nanosecond-scale prediction speeds while reducing the required resources as compares to traditional binary neural networks. Despite these benefits, the trade-offs between LGN parameters and resulting hardware synthesis characteristics are not well characterized. This paper therefore studies the tradeoffs between power, resource utilization, inference speed, and model accuracy when varying the depth and width of LGNs synthesized for Field Programmable Gate Arrays (FPGA). Results reveal that the final layer of an LGN is critical to minimize timing and resource usage (i.e. 28\% decrease), as this layer dictates the logic size of summing operations. Subject to timing and routing constraints, deeper and wider LGNs can be synthesized for FPGA when the final layer is narrow. Further tradeoffs are presented to help ML engineers select baseline LGN architectures for FPGAs with a set number of Look Up Tables (LUT).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。