arXiv:2502.02304hep-excs.DC2025-02被引 7

对比FPGA与GPU在粒子追踪中的推理性能,发现FPGA更省电且延迟更低。

Comparative Analysis of FPGA and GPU Performance for Machine Learning-Based Track Reconstruction at LHCb

  • 用HLS4ML将多层感知机部署到FPGA,实现高效推理
  • FPGA吞吐量达12.5万事件/秒,比GPU快1.8倍
  • 适合对功耗和延迟敏感的高能物理实时处理场景

在高能物理领域,大型强子对撞机(LHC)的亮度和探测器精细度不断提升,对数据处理效率提出更高要求。机器学习因其计算复杂度与探测器事例数呈线性关系,成为重建带电粒子轨迹的有力工具。目前,基于图神经网络的轨迹重建流水线已在LHCb实验的初筛触发系统中部署于GPU,为架构比较提供了基础。本文针对该流程的第一步——多层感知机实现,首次系统比较了FPGA与GPU在机器学习推理方面的性能表现。通过HLS4ML工具链部署FPGA方案,与GPU实现进行基准测试。结果表明,FPGA在保持低延迟的同时,达到12.5万事件/秒的吞吐量,相较GPU提升1.8倍,并显著降低功耗。该研究展示了无需专业FPGA开发经验即可实现高性能推理的可行性,凸显了FPGA在高通量、低延迟场景中的应用潜力。

原文摘要 · Abstract (English)

In high-energy physics, the increasing luminosity and detector granularity at the Large Hadron Collider are driving the need for more efficient data processing solutions. Machine Learning has emerged as a promising tool for reconstructing charged particle tracks, due to its potentially linear computational scaling with detector hits. The recent implementation of a graph neural network-based track reconstruction pipeline in the first level trigger of the LHCb experiment on GPUs serves as a platform for comparative studies between computational architectures in the context of high-energy physics. This paper presents a novel comparison of the throughput of ML model inference between FPGAs and GPUs, focusing on the first step of the track reconstruction pipeline$\unicode{x2013}$an implementation of a multilayer perceptron. Using HLS4ML for FPGA deployment, we benchmark its performance against the GPU implementation and demonstrate the potential of FPGAs for high-throughput, low-latency inference without the need for an expertise in FPGA development and while consuming significantly less power.

FPGA机器学习粒子物理低延迟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。