arXiv:2409.05207cs.LG2024-09被引 11

在FPGA上实现低延迟Transformer推理,支持物理实时分析。

Low Latency Transformer Inference on FPGAs for Physics Applications with hls4ml

  • 用hls4ml在FPGA上部署多头注意力等核心模块
  • 三类模型在VU13P芯片上延迟均低于2微秒
  • 兼容任意TensorFlow构建的Transformer,适合高能物理应用

本研究利用hls4ml在现场可编程门阵列(FPGA)上高效实现Transformer架构。我们提出了多头注意力、Softmax和归一化层的实现策略,并评估了三种不同模型。这些模型在VU13P FPGA芯片上的部署延迟均低于2微秒,展示了其在实时应用中的潜力。hls4ml对任意基于TensorFlow构建的Transformer模型的兼容性,进一步提升了该工作的可扩展性与适用范围。

原文摘要 · Abstract (English)

This study presents an efficient implementation of transformer architectures in Field-Programmable Gate Arrays(FPGAs) using hls4ml. We demonstrate the strategy for implementing the multi-head attention, softmax, and normalization layer and evaluate three distinct models. Their deployment on VU13P FPGA chip achieved latency less than 2us, demonstrating the potential for real-time applications. HLS4ML compatibility with any TensorFlow-built transformer model further enhances the scalability and applicability of this work. Index Terms: FPGAs, machine learning, transformers, high energy physics, LIGO

FPGATransformer低延迟高能物理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。