arXiv:2512.01463cs.ARcs.LG2025-12被引 18

将深度学习模型转为FPGA/ASIC可部署代码,实现低延迟高效推理。

hls4ml: A Flexible, Open-Source Platform for Deep Learning Acceleration on Reconfigurable Hardware

  • 支持多种框架与编译器的灵活转换,打通软硬件协同设计链路。
  • 可生成低延迟、低功耗的HLS代码,适配科学与商业场景需求。
  • 开源平台助力高能物理等领域的实时推理加速,适合硬件工程师使用。

我们提出 hls4ml,一个免费开源平台,可将现代深度学习框架中的机器学习模型转换为可用于现场可编程门阵列(FPGAs)或专用集成电路(ASICs)的高层次综合(HLS)代码。其灵活模块化设计支持多种深度学习框架,并兼容 Vitis HLS、Intel oneAPI 及 Catapult HLS 等多个厂商的HLS编译器。结合完整的软硬件协同设计生态,hls4ml 已在低延迟、低资源占用和低功耗至关重要的商业与科学应用中实现机器学习推理加速。本文详细描述了 hls4ml 的架构与功能,讨论了生成HLS代码的核心设计原则,并展示了若干性能结果。

原文摘要 · Abstract (English)

We present hls4ml, a free and open-source platform that translates machine learning (ML) models from modern deep learning frameworks into high-level synthesis (HLS) code that can be integrated into full designs for field-programmable gate arrays (FPGAs) or application-specific integrated circuits (ASICs). With its flexible and modular design, hls4ml supports a large number of deep learning frameworks and can target HLS compilers from several vendors, including Vitis HLS, Intel oneAPI and Catapult HLS. Together with a wider eco-system for software-hardware co-design, hls4ml has enabled the acceleration of ML inference in a wide range of commercial and scientific applications where low latency, resource usage, and power consumption are critical. In this paper, we describe the structure and functionality of the hls4ml platform. The overarching design considerations for the generated HLS code are discussed, together with selected performance results.

硬件加速FPGA深度学习部署开源工具

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。